{
  "id": 60455,
  "title": "Experimental LSTM approach",
  "url": "/competitions/trackml-particle-identification/discussion/60455",
  "author_name": "",
  "post_date": "2018-07-04T19:30:36.270876900Z",
  "votes": 13,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I thought LSTM was going to be the right approach at the very beginning of this competition and finally got a chance to try it after being inspired by Heng's post - <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/59154\">https://www.kaggle.com/c/trackml-particle-identification/discussion/59154</a></p>\n\n<ul>\n<li>The result was trained on 1700 events only on the first dimension when <code>x,y,z &gt; 0</code>  - I have tried different dimensions but the predicted results were not even close, so I'd have to train more models for different dimensions. </li>\n<li>The input features were normalized <code>phi, r, z</code> or <code>x, y, z</code> and they yielded similar performance.  It's a bit more accurate with <code>phi, r, z</code></li>\n<li>It uses 5 hits to predict coordinates of next 5 hits (input dimension = <code>10x3</code> - 5 hits were padded with <code>0</code>, output dimension = <code>10x3</code> as well)</li>\n<li>See the model code below, I tried different numbers of neurons, <code>24, 48 and 96</code>, they don't make a big difference in performance</li>\n<li>It was good at predicting tendency but it didn't learn to predict precise z coordinates, esp. when there are two very close hits with <code>z1-z2 = 0.5mm</code> (two hits in joint modules in the same layer, see the 10th slide - <a href=\"https://www.nikhef.nl/pub/conferences/Vertex2009/talks/Migliore_vertex2009.pdf\">https://www.nikhef.nl/pub/conferences/Vertex2009/talks/Migliore_vertex2009.pdf</a>)</li>\n<li><p>When I used dbscan-predicted tracks as seeds, the prediction was way off, the reason was that my predicted dbscan tracks only get <code>&lt; 90%</code> of seeds.</p></li>\n<li><p>However, it has its potential to be used for outlier removal and track extension when we train an LSTM only with filtered data. </p>\n\n<pre><code>   def build_model(num_hidden, input_shape, output_shape, loss='mse', optimizer='Nadam'):\n\n       inputs = layers.Input(shape=input_shape)\n       hidden = layers.LSTM(units=num_hidden, return_sequences=True)(inputs)\n       outputs = layers.TimeDistributed(layers.Dense(output_shape[1], activation='linear') (hidden)\n       model = models.Model(inputs=inputs, outputs=outputs)\n       if GPU &gt; 0:\n           gpu_model = multi_gpu_model(model, GPU)\n       else:\n           gpu_model = model\n\n           gpu_model.compile(loss=loss, optimizer=optimizer)\n\n       return model, gpu_model\n</code></pre></li>\n<li><p>Thanks to @Heng's beautiful visualization code, the first image shows predicted tracks in light-grey and the ground truth in colour \n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/352645/9782/event_1003_LSTM_5_seeded_hits.png\" alt=\"ground truth vs predicted tracks\"></p></li>\n<li>The second image shows using predicted tracks using 5-hit-seed predicted by dbscan in light-grey and the predicted dbscan full tracks in colour\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/352645/9783/event_1003_dbscan_LSTM_5_seeded_hits.png\" alt=\"dbscan tracks vs predicted tracks\"></li>\n</ul>",
  "messages": [
    {
      "id": "352645",
      "postDate": "07/04/2018 19:30:36",
      "content": "<p>I thought LSTM was going to be the right approach at the very beginning of this competition and finally got a chance to try it after being inspired by Heng's post - <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/59154\">https://www.kaggle.com/c/trackml-particle-identification/discussion/59154</a></p>\n\n<ul>\n<li>The result was trained on 1700 events only on the first dimension when <code>x,y,z &gt; 0</code>  - I have tried different dimensions but the predicted results were not even close, so I'd have to train more models for different dimensions. </li>\n<li>The input features were normalized <code>phi, r, z</code> or <code>x, y, z</code> and they yielded similar performance.  It's a bit more accurate with <code>phi, r, z</code></li>\n<li>It uses 5 hits to predict coordinates of next 5 hits (input dimension = <code>10x3</code> - 5 hits were padded with <code>0</code>, output dimension = <code>10x3</code> as well)</li>\n<li>See the model code below, I tried different numbers of neurons, <code>24, 48 and 96</code>, they don't make a big difference in performance</li>\n<li>It was good at predicting tendency but it didn't learn to predict precise z coordinates, esp. when there are two very close hits with <code>z1-z2 = 0.5mm</code> (two hits in joint modules in the same layer, see the 10th slide - <a href=\"https://www.nikhef.nl/pub/conferences/Vertex2009/talks/Migliore_vertex2009.pdf\">https://www.nikhef.nl/pub/conferences/Vertex2009/talks/Migliore_vertex2009.pdf</a>)</li>\n<li><p>When I used dbscan-predicted tracks as seeds, the prediction was way off, the reason was that my predicted dbscan tracks only get <code>&lt; 90%</code> of seeds.</p></li>\n<li><p>However, it has its potential to be used for outlier removal and track extension when we train an LSTM only with filtered data. </p>\n\n<pre><code>   def build_model(num_hidden, input_shape, output_shape, loss='mse', optimizer='Nadam'):\n\n       inputs = layers.Input(shape=input_shape)\n       hidden = layers.LSTM(units=num_hidden, return_sequences=True)(inputs)\n       outputs = layers.TimeDistributed(layers.Dense(output_shape[1], activation='linear') (hidden)\n       model = models.Model(inputs=inputs, outputs=outputs)\n       if GPU &gt; 0:\n           gpu_model = multi_gpu_model(model, GPU)\n       else:\n           gpu_model = model\n\n           gpu_model.compile(loss=loss, optimizer=optimizer)\n\n       return model, gpu_model\n</code></pre></li>\n<li><p>Thanks to @Heng's beautiful visualization code, the first image shows predicted tracks in light-grey and the ground truth in colour \n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/352645/9782/event_1003_LSTM_5_seeded_hits.png\" alt=\"ground truth vs predicted tracks\"></p></li>\n<li>The second image shows using predicted tracks using 5-hit-seed predicted by dbscan in light-grey and the predicted dbscan full tracks in colour\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/352645/9783/event_1003_dbscan_LSTM_5_seeded_hits.png\" alt=\"dbscan tracks vs predicted tracks\"></li>\n</ul>",
      "rawMarkdown": "I thought LSTM was going to be the right approach at the very beginning of this competition and finally got a chance to try it after being inspired by Heng's post - https://www.kaggle.com/c/trackml-particle-identification/discussion/59154\n\n - The result was trained on 1700 events only on the first dimension when `x,y,z &gt; 0`  - I have tried different dimensions but the predicted results were not even close, so I'd have to train more models for different dimensions. \n - The input features were normalized `phi, r, z` or `x, y, z` and they yielded similar performance.  It's a bit more accurate with `phi, r, z`\n - It uses 5 hits to predict coordinates of next 5 hits (input dimension = `10x3` - 5 hits were padded with `0`, output dimension = `10x3` as well)\n - See the model code below, I tried different numbers of neurons, `24, 48 and 96`, they don't make a big difference in performance\n - It was good at predicting tendency but it didn't learn to predict precise z coordinates, esp. when there are two very close hits with `z1-z2 = 0.5mm` (two hits in joint modules in the same layer, see the 10th slide - https://www.nikhef.nl/pub/conferences/Vertex2009/talks/Migliore_vertex2009.pdf)\n - When I used dbscan-predicted tracks as seeds, the prediction was way off, the reason was that my predicted dbscan tracks only get `&lt; 90%` of seeds.\n\n - However, it has its potential to be used for outlier removal and track extension when we train an LSTM only with filtered data. \n\n           def build_model(num_hidden, input_shape, output_shape, loss='mse', optimizer='Nadam'):\n            \n               inputs = layers.Input(shape=input_shape)\n               hidden = layers.LSTM(units=num_hidden, return_sequences=True)(inputs)\n               outputs = layers.TimeDistributed(layers.Dense(output_shape[1], activation='linear') (hidden)\n               model = models.Model(inputs=inputs, outputs=outputs)\n               if GPU &gt; 0:\n                   gpu_model = multi_gpu_model(model, GPU)\n               else:\n                   gpu_model = model\n      \n                   gpu_model.compile(loss=loss, optimizer=optimizer)\n         \n               return model, gpu_model\n     \n - Thanks to @Heng's beautiful visualization code, the first image shows predicted tracks in light-grey and the ground truth in colour \n![ground truth vs predicted tracks][1]\n - The second image shows using predicted tracks using 5-hit-seed predicted by dbscan in light-grey and the predicted dbscan full tracks in colour\n ![dbscan tracks vs predicted tracks][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/352645/9782/event_1003_LSTM_5_seeded_hits.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/352645/9783/event_1003_dbscan_LSTM_5_seeded_hits.png",
      "votes": null
    },
    {
      "id": "352777",
      "postDate": "07/05/2018 05:12:31",
      "content": "<p>i notice this too. I suggest instead of predicting the (x,yz) directly , you can try the below suggestions. You can change (x,yz) into other representations like (a,r,z),  or (x,y,z,a,r)  etc  ...</p>\n\n<p>0) this doesn't work:    hidden,cell,(x_t,y_t,z_t) ---&gt; [LSTM] ---&gt; predicted (x_t+1,y_t+1,z_t+1)</p>\n\n<p>1) suggestion 1: hidden,cell,(z) ---&gt; [LSTM] ---&gt; predicted (x,y)</p>\n\n<p>2) suggestion 2: hidden,cell,( nearest neighbour in next z, n1,n2,n3,n4 .....nN) ---&gt; [LSTM] ---&gt; predicted n link probabilites.</p>\n\n<p>n1 =(x,y,z)= some neighbour hit candidate from KNN, etc</p>\n\n<p>in suggestion (1) and (2), z is taken as the \"time\" of the sequence</p>\n\n<p>3) hidden,cell,(x,y,z) ---&gt; [LSTM] ---&gt; predicted (helix curve parameters, ... or some parametric curve). The helix equation is not exact. You can use other parametric curve formula.</p>",
      "rawMarkdown": "i notice this too. I suggest instead of predicting the (x,yz) directly , you can try the below suggestions. You can change (x,yz) into other representations like (a,r,z),  or (x,y,z,a,r)  etc  ...\n\n0) this doesn't work:    hidden,cell,(x_t,y_t,z_t) ---&gt; [LSTM] ---&gt; predicted (x_t+1,y_t+1,z_t+1)\n\n1) suggestion 1: hidden,cell,(z) ---&gt; [LSTM] ---&gt; predicted (x,y)\n\n2) suggestion 2: hidden,cell,( nearest neighbour in next z, n1,n2,n3,n4 .....nN) ---&gt; [LSTM] ---&gt; predicted n link probabilites.\n\nn1 =(x,y,z)= some neighbour hit candidate from KNN, etc\n\n\nin suggestion (1) and (2), z is taken as the \"time\" of the sequence\n\n\n3) hidden,cell,(x,y,z) ---&gt; [LSTM] ---&gt; predicted (helix curve parameters, ... or some parametric curve). The helix equation is not exact. You can use other parametric curve formula.",
      "votes": null
    },
    {
      "id": "352833",
      "postDate": "07/05/2018 08:13:16",
      "content": "<p>I'm always wondering why one would want to use a RNN when having fixed length sequences.  Why not use a feed forward network to predict 5 hits coordinates from 5 seed hits?</p>",
      "rawMarkdown": "I'm always wondering why one would want to use a RNN when having fixed length sequences.  Why not use a feed forward network to predict 5 hits coordinates from 5 seed hits?",
      "votes": null
    },
    {
      "id": "352861",
      "postDate": "07/05/2018 09:50:49",
      "content": "<p>Thanks @CPMP for your input, originally I was hoping an LSTM could find the correlation between the first few hits and the final hits  by using its traits such as memory cells and update/forget gates, even when there are no temporal traits in the hits since all hits were measured at almost the same time at light-speed, we could still map detector information or z-axis to time steps. But I think you're right a simple feedfoward NN can do a similar job too if it's not better. The problem I see with an RNN model is that we need very precise predictions within a very small margin error, not like a traditional sales prediction that doesn't have the same requirement. An NN probably wouldn't give us precise enough predictions, a brute force search would do a way better job to find the trend, we'd need to do lots of post processing using an NN. However, since <code>CERN</code> was interested in replacing the Kalman filter with LSTM due to Kalman filter's time complexity(?), I still wanted to give it a shot. </p>",
      "rawMarkdown": "Thanks @CPMP for your input, originally I was hoping an LSTM could find the correlation between the first few hits and the final hits  by using its traits such as memory cells and update/forget gates, even when there are no temporal traits in the hits since all hits were measured at almost the same time at light-speed, we could still map detector information or z-axis to time steps. But I think you're right a simple feedfoward NN can do a similar job too if it's not better. The problem I see with an RNN model is that we need very precise predictions within a very small margin error, not like a traditional sales prediction that doesn't have the same requirement. An NN probably wouldn't give us precise enough predictions, a brute force search would do a way better job to find the trend, we'd need to do lots of post processing using an NN. However, since `CERN` was interested in replacing the Kalman filter with LSTM due to Kalman filter's time complexity(?), I still wanted to give it a shot.",
      "votes": null
    },
    {
      "id": "352863",
      "postDate": "07/05/2018 09:54:32",
      "content": "<p>@Heng, thanks for your suggestions, I'll try some of them for sure. Have you solved the angular discontinuity problem and the cross-dimension problem (x,y,z are not all in the first dimension)? With this approach, we'd probably miss cross-dimensional tracks.  </p>",
      "rawMarkdown": "Heng, thanks for your suggestions, I'll try some of them for sure. Have you solved the angular discontinuity problem and the cross-dimension problem (x,y,z are not all in the first dimension)? With this approach, we'd probably miss cross-dimensional tracks.",
      "votes": null
    },
    {
      "id": "352876",
      "postDate": "07/05/2018 10:29:13",
      "content": "<p>if you are using a,r,z coordinate, you can repeat your array:</p>\n\n<p>e.g. original input</p>\n\n<pre><code>1 2 3 4 5\n0 7 7 8 9\n4 5 8 0 2\n</code></pre>\n\n<p>after repeating</p>\n\n<pre><code>... 4 5    1 2 3 4 5   1 2... \n... 8 9    0 7 7 8 9   0 7... \n... 0 2    4 5 8 0 2   4 5... \n</code></pre>\n\n<p>the angle dimension is warped and repeated</p>",
      "rawMarkdown": "if you are using a,r,z coordinate, you can repeat your array:\n\ne.g. original input\n\n    1 2 3 4 5\n    0 7 7 8 9\n    4 5 8 0 2\n\nafter repeating\n\n    ... 4 5    1 2 3 4 5   1 2... \n    ... 8 9    0 7 7 8 9   0 7... \n    ... 0 2    4 5 8 0 2   4 5... \n     \nthe angle dimension is warped and repeated",
      "votes": null
    },
    {
      "id": "352882",
      "postDate": "07/05/2018 10:38:43",
      "content": "<p>the reason why LSTM is not accurate could be your mse loss.</p>\n\n<p>Assume in the training data, there are 2 sample sequence ((0,1,2), (0,2,2)) and  ((0,1,2), (0,2,5))</p>\n\n<p>for the same input (0,1,2), the predicted output to minimize mse loss is mean( (0,2,2),(0,2,5) ). </p>\n\n<hr>\n\n<p>on a side note, i did an experiment using RANSAC. First i verify that my parametric form is correct. I use non-linear least square to recover the parameters and the parametric curve can fit the ground truth track well.</p>\n\n<p>then for a given small volume of hits, i can find several curves to fit the hits with  low loss. These curves include true curves and outliers. </p>",
      "rawMarkdown": "the reason why LSTM is not accurate could be your mse loss.\n\nAssume in the training data, there are 2 sample sequence ((0,1,2), (0,2,2)) and  ((0,1,2), (0,2,5))\n\nfor the same input (0,1,2), the predicted output to minimize mse loss is mean( (0,2,2),(0,2,5) ). \n\n---\n\non a side note, i did an experiment using RANSAC. First i verify that my parametric form is correct. I use non-linear least square to recover the parameters and the parametric curve can fit the ground truth track well.\n\nthen for a given small volume of hits, i can find several curves to fit the hits with  low loss. These curves include true curves and outliers.",
      "votes": null
    },
    {
      "id": "352886",
      "postDate": "07/05/2018 10:54:56",
      "content": "<p>@Heng, interesting, I thought about training a model to predict parameter space but not all tracks follow the helix parameter space and least square probably would do a better job than an NN. Yeah, I need to rewrite the loss function instead of using the default Keras mse, my loss didn't really go down after 10 epochs or so and I noticed this problem too. </p>",
      "rawMarkdown": "Heng, interesting, I thought about training a model to predict parameter space but not all tracks follow the helix parameter space and least square probably would do a better job than an NN. Yeah, I need to rewrite the loss function instead of using the default Keras mse, my loss didn't really go down after 10 epochs or so and I noticed this problem too.",
      "votes": null
    },
    {
      "id": "353342",
      "postDate": "07/06/2018 13:36:30",
      "content": "<p>@Nicole Finnie</p>\n\n<p>an example of predicting pairwise link can be found at:</p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60447\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60447</a></p>",
      "rawMarkdown": "Nicole Finnie\n\nan example of predicting pairwise link can be found at:\n\nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/60447",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 352777,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/05/2018 05:12:31",
      "content": "<p>i notice this too. I suggest instead of predicting the (x,yz) directly , you can try the below suggestions. You can change (x,yz) into other representations like (a,r,z),  or (x,y,z,a,r)  etc  ...</p>\n\n<p>0) this doesn't work:    hidden,cell,(x_t,y_t,z_t) ---&gt; [LSTM] ---&gt; predicted (x_t+1,y_t+1,z_t+1)</p>\n\n<p>1) suggestion 1: hidden,cell,(z) ---&gt; [LSTM] ---&gt; predicted (x,y)</p>\n\n<p>2) suggestion 2: hidden,cell,( nearest neighbour in next z, n1,n2,n3,n4 .....nN) ---&gt; [LSTM] ---&gt; predicted n link probabilites.</p>\n\n<p>n1 =(x,y,z)= some neighbour hit candidate from KNN, etc</p>\n\n<p>in suggestion (1) and (2), z is taken as the \"time\" of the sequence</p>\n\n<p>3) hidden,cell,(x,y,z) ---&gt; [LSTM] ---&gt; predicted (helix curve parameters, ... or some parametric curve). The helix equation is not exact. You can use other parametric curve formula.</p>",
      "votes": null,
      "replies": [
        {
          "id": 352863,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "07/05/2018 09:54:32",
          "content": "<p>@Heng, thanks for your suggestions, I'll try some of them for sure. Have you solved the angular discontinuity problem and the cross-dimension problem (x,y,z are not all in the first dimension)? With this approach, we'd probably miss cross-dimensional tracks.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 352876,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/05/2018 10:29:13",
          "content": "<p>if you are using a,r,z coordinate, you can repeat your array:</p>\n\n<p>e.g. original input</p>\n\n<pre><code>1 2 3 4 5\n0 7 7 8 9\n4 5 8 0 2\n</code></pre>\n\n<p>after repeating</p>\n\n<pre><code>... 4 5    1 2 3 4 5   1 2... \n... 8 9    0 7 7 8 9   0 7... \n... 0 2    4 5 8 0 2   4 5... \n</code></pre>\n\n<p>the angle dimension is warped and repeated</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 352833,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/05/2018 08:13:16",
      "content": "<p>I'm always wondering why one would want to use a RNN when having fixed length sequences.  Why not use a feed forward network to predict 5 hits coordinates from 5 seed hits?</p>",
      "votes": null,
      "replies": [
        {
          "id": 352861,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "07/05/2018 09:50:49",
          "content": "<p>Thanks @CPMP for your input, originally I was hoping an LSTM could find the correlation between the first few hits and the final hits  by using its traits such as memory cells and update/forget gates, even when there are no temporal traits in the hits since all hits were measured at almost the same time at light-speed, we could still map detector information or z-axis to time steps. But I think you're right a simple feedfoward NN can do a similar job too if it's not better. The problem I see with an RNN model is that we need very precise predictions within a very small margin error, not like a traditional sales prediction that doesn't have the same requirement. An NN probably wouldn't give us precise enough predictions, a brute force search would do a way better job to find the trend, we'd need to do lots of post processing using an NN. However, since <code>CERN</code> was interested in replacing the Kalman filter with LSTM due to Kalman filter's time complexity(?), I still wanted to give it a shot. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 352882,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/05/2018 10:38:43",
      "content": "<p>the reason why LSTM is not accurate could be your mse loss.</p>\n\n<p>Assume in the training data, there are 2 sample sequence ((0,1,2), (0,2,2)) and  ((0,1,2), (0,2,5))</p>\n\n<p>for the same input (0,1,2), the predicted output to minimize mse loss is mean( (0,2,2),(0,2,5) ). </p>\n\n<hr>\n\n<p>on a side note, i did an experiment using RANSAC. First i verify that my parametric form is correct. I use non-linear least square to recover the parameters and the parametric curve can fit the ground truth track well.</p>\n\n<p>then for a given small volume of hits, i can find several curves to fit the hits with  low loss. These curves include true curves and outliers. </p>",
      "votes": null,
      "replies": [
        {
          "id": 352886,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "07/05/2018 10:54:56",
          "content": "<p>@Heng, interesting, I thought about training a model to predict parameter space but not all tracks follow the helix parameter space and least square probably would do a better job than an NN. Yeah, I need to rewrite the loss function instead of using the default Keras mse, my loss didn't really go down after 10 epochs or so and I noticed this problem too. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 353342,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/06/2018 13:36:30",
          "content": "<p>@Nicole Finnie</p>\n\n<p>an example of predicting pairwise link can be found at:</p>\n\n<p><a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60447\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60447</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "352645": "I thought LSTM was going to be the right approach at the very beginning of this competition and finally got a chance to try it after being inspired by Heng's post - https://www.kaggle.com/c/trackml-particle-identification/discussion/59154\n\n - The result was trained on 1700 events only on the first dimension when `x,y,z &gt; 0`  - I have tried different dimensions but the predicted results were not even close, so I'd have to train more models for different dimensions. \n - The input features were normalized `phi, r, z` or `x, y, z` and they yielded similar performance.  It's a bit more accurate with `phi, r, z`\n - It uses 5 hits to predict coordinates of next 5 hits (input dimension = `10x3` - 5 hits were padded with `0`, output dimension = `10x3` as well)\n - See the model code below, I tried different numbers of neurons, `24, 48 and 96`, they don't make a big difference in performance\n - It was good at predicting tendency but it didn't learn to predict precise z coordinates, esp. when there are two very close hits with `z1-z2 = 0.5mm` (two hits in joint modules in the same layer, see the 10th slide - https://www.nikhef.nl/pub/conferences/Vertex2009/talks/Migliore_vertex2009.pdf)\n - When I used dbscan-predicted tracks as seeds, the prediction was way off, the reason was that my predicted dbscan tracks only get `&lt; 90%` of seeds.\n\n - However, it has its potential to be used for outlier removal and track extension when we train an LSTM only with filtered data. \n\n           def build_model(num_hidden, input_shape, output_shape, loss='mse', optimizer='Nadam'):\n            \n               inputs = layers.Input(shape=input_shape)\n               hidden = layers.LSTM(units=num_hidden, return_sequences=True)(inputs)\n               outputs = layers.TimeDistributed(layers.Dense(output_shape[1], activation='linear') (hidden)\n               model = models.Model(inputs=inputs, outputs=outputs)\n               if GPU &gt; 0:\n                   gpu_model = multi_gpu_model(model, GPU)\n               else:\n                   gpu_model = model\n      \n                   gpu_model.compile(loss=loss, optimizer=optimizer)\n         \n               return model, gpu_model\n     \n - Thanks to @Heng's beautiful visualization code, the first image shows predicted tracks in light-grey and the ground truth in colour \n![ground truth vs predicted tracks][1]\n - The second image shows using predicted tracks using 5-hit-seed predicted by dbscan in light-grey and the predicted dbscan full tracks in colour\n ![dbscan tracks vs predicted tracks][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/352645/9782/event_1003_LSTM_5_seeded_hits.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/352645/9783/event_1003_dbscan_LSTM_5_seeded_hits.png",
    "352777": "i notice this too. I suggest instead of predicting the (x,yz) directly , you can try the below suggestions. You can change (x,yz) into other representations like (a,r,z),  or (x,y,z,a,r)  etc  ...\n\n0) this doesn't work:    hidden,cell,(x_t,y_t,z_t) ---&gt; [LSTM] ---&gt; predicted (x_t+1,y_t+1,z_t+1)\n\n1) suggestion 1: hidden,cell,(z) ---&gt; [LSTM] ---&gt; predicted (x,y)\n\n2) suggestion 2: hidden,cell,( nearest neighbour in next z, n1,n2,n3,n4 .....nN) ---&gt; [LSTM] ---&gt; predicted n link probabilites.\n\nn1 =(x,y,z)= some neighbour hit candidate from KNN, etc\n\n\nin suggestion (1) and (2), z is taken as the \"time\" of the sequence\n\n\n3) hidden,cell,(x,y,z) ---&gt; [LSTM] ---&gt; predicted (helix curve parameters, ... or some parametric curve). The helix equation is not exact. You can use other parametric curve formula.",
    "352833": "I'm always wondering why one would want to use a RNN when having fixed length sequences.  Why not use a feed forward network to predict 5 hits coordinates from 5 seed hits?",
    "352861": "Thanks @CPMP for your input, originally I was hoping an LSTM could find the correlation between the first few hits and the final hits  by using its traits such as memory cells and update/forget gates, even when there are no temporal traits in the hits since all hits were measured at almost the same time at light-speed, we could still map detector information or z-axis to time steps. But I think you're right a simple feedfoward NN can do a similar job too if it's not better. The problem I see with an RNN model is that we need very precise predictions within a very small margin error, not like a traditional sales prediction that doesn't have the same requirement. An NN probably wouldn't give us precise enough predictions, a brute force search would do a way better job to find the trend, we'd need to do lots of post processing using an NN. However, since `CERN` was interested in replacing the Kalman filter with LSTM due to Kalman filter's time complexity(?), I still wanted to give it a shot.",
    "352863": "Heng, thanks for your suggestions, I'll try some of them for sure. Have you solved the angular discontinuity problem and the cross-dimension problem (x,y,z are not all in the first dimension)? With this approach, we'd probably miss cross-dimensional tracks.",
    "352876": "if you are using a,r,z coordinate, you can repeat your array:\n\ne.g. original input\n\n    1 2 3 4 5\n    0 7 7 8 9\n    4 5 8 0 2\n\nafter repeating\n\n    ... 4 5    1 2 3 4 5   1 2... \n    ... 8 9    0 7 7 8 9   0 7... \n    ... 0 2    4 5 8 0 2   4 5... \n     \nthe angle dimension is warped and repeated",
    "352882": "the reason why LSTM is not accurate could be your mse loss.\n\nAssume in the training data, there are 2 sample sequence ((0,1,2), (0,2,2)) and  ((0,1,2), (0,2,5))\n\nfor the same input (0,1,2), the predicted output to minimize mse loss is mean( (0,2,2),(0,2,5) ). \n\n---\n\non a side note, i did an experiment using RANSAC. First i verify that my parametric form is correct. I use non-linear least square to recover the parameters and the parametric curve can fit the ground truth track well.\n\nthen for a given small volume of hits, i can find several curves to fit the hits with  low loss. These curves include true curves and outliers.",
    "352886": "Heng, interesting, I thought about training a model to predict parameter space but not all tracks follow the helix parameter space and least square probably would do a better job than an NN. Yeah, I need to rewrite the loss function instead of using the default Keras mse, my loss didn't really go down after 10 epochs or so and I noticed this problem too.",
    "353342": "Nicole Finnie\n\nan example of predicting pairwise link can be found at:\n\nhttps://www.kaggle.com/c/trackml-particle-identification/discussion/60447"
  },
  "source": "meta"
}