{
  "id": 59154,
  "title": "anyone start using lstm yet?",
  "url": "/competitions/trackml-particle-identification/discussion/59154",
  "author_name": "",
  "post_date": "2018-06-19T07:48:15.899033700Z",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>has anyone start using lstm yet? lstm seems to be a potential candidate to replace kalman filter. Anyone has any results with lstm?</p>",
  "messages": [
    {
      "id": "345100",
      "postDate": "06/19/2018 07:48:15",
      "content": "<p>has anyone start using lstm yet? lstm seems to be a potential candidate to replace kalman filter. Anyone has any results with lstm?</p>",
      "rawMarkdown": "has anyone start using lstm yet? lstm seems to be a potential candidate to replace kalman filter. Anyone has any results with lstm?",
      "votes": null
    },
    {
      "id": "345134",
      "postDate": "06/19/2018 08:47:03",
      "content": "<p>@Heng, I encountered problems when designing an LSTM so I stopped. Some coordinates have almost the same directions when their radius and angular distance are almost the same. Sometimes both of them belong to the same track (one may be a reconstructed hit), and sometimes they belong to different tracks.  Do you think an LSTM can tell them apart? I don't see how it can be more accurate, it would require heavy post processing just like the current clustering approach does. Has any of your DL models outperformed your clustering result?</p>",
      "rawMarkdown": "Heng, I encountered problems when designing an LSTM so I stopped. Some coordinates have almost the same directions when their radius and angular distance are almost the same. Sometimes both of them belong to the same track (one may be a reconstructed hit), and sometimes they belong to different tracks.  Do you think an LSTM can tell them apart? I don't see how it can be more accurate, it would require heavy post processing just like the current clustering approach does. Has any of your DL models outperformed your clustering result?",
      "votes": null
    },
    {
      "id": "345216",
      "postDate": "06/19/2018 12:18:19",
      "content": "<p>i haven't study lstm in detail yet. this is my results on test (not train):</p>\n\n<p>left:  x,y,z\nright: r,angle,z</p>\n\n<p>gray: ground truth.</p>\n\n<p>color: lstm results</p>\n\n<p>reference: \n<a href=\"https://indico.cern.ch/event/658267/contributions/2881175/attachments/1621912/2581064/Farrell_heptrkx_ctd2018.pdf\">https://indico.cern.ch/event/658267/contributions/2881175/attachments/1621912/2581064/Farrell_heptrkx_ctd2018.pdf</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/345216/9641/LSTM_test.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "i haven't study lstm in detail yet. this is my results on test (not train):\n\nleft:  x,y,z\nright: r,angle,z\n\ngray: ground truth.\n\ncolor: lstm results\n\nreference: \nhttps://indico.cern.ch/event/658267/contributions/2881175/attachments/1621912/2581064/Farrell_heptrkx_ctd2018.pdf\n\n\n ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/345216/9641/LSTM_test.png",
      "votes": null
    },
    {
      "id": "345514",
      "postDate": "06/20/2018 01:53:46",
      "content": "<p>i did some experiments. learning based lstm is better than my track extension code.</p>\n\n<p>(and i think learning-based ensemble method will work better than hand crafted ones too, and i am working on that too)</p>",
      "rawMarkdown": "i did some experiments. learning based lstm is better than my track extension code.\n\n(and i think learning-based ensemble method will work better than hand crafted ones too, and i am working on that too)",
      "votes": null
    },
    {
      "id": "346774",
      "postDate": "06/22/2018 11:00:53",
      "content": "<p>Thanks for sharing, improving clustering postprocessing is also what I am looking at now.  Looks like the next low hanging fruit. </p>",
      "rawMarkdown": "Thanks for sharing, improving clustering postprocessing is also what I am looking at now.  Looks like the next low hanging fruit.",
      "votes": null
    },
    {
      "id": "346809",
      "postDate": "06/22/2018 13:22:50",
      "content": "<p>using dbscan for seeding tracks</p>",
      "rawMarkdown": "using dbscan for seeding tracks",
      "votes": null
    },
    {
      "id": "346852",
      "postDate": "06/22/2018 15:17:37",
      "content": "<p>updated code to analyse missed and false positive seeding tracks</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/346852/9666/error2.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/346852/9667/error1.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "updated code to analyse missed and false positive seeding tracks\n\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/346852/9666/error2.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/346852/9667/error1.png",
      "votes": null
    },
    {
      "id": "347312",
      "postDate": "06/23/2018 22:03:36",
      "content": "<p>In the script provided, you seed hits from the region z &gt; 500 and r &lt; 50, right? This captures a portion of detectors in volume 7 and 9. Why do you only seed hits from the end caps and not the 'barrel'-type detectors (use no z constraint)?</p>",
      "rawMarkdown": "In the script provided, you seed hits from the region z &gt; 500 and r &lt; 50, right? This captures a portion of detectors in volume 7 and 9. Why do you only seed hits from the end caps and not the 'barrel'-type detectors (use no z constraint)?",
      "votes": null
    },
    {
      "id": "347597",
      "postDate": "06/24/2018 19:56:46",
      "content": "<p>@Matthew the portion is in the volume 7, for volume 9, it should be z &lt; -500. And my approximate scoring on the volume 8 is <code>np.absolute(z) &lt; 500</code>. </p>\n\n<p>@Heng the portion scoring function is a bit optimistic since it doesn't consider the factor of \"losing majority\" of the extended track but merely takes the portion of detected hits from the dbscan cluster and scores on it.  Please correct me if I was wrong. However, it serves its purpose well and gave us a very good insight. Thanks for letting me steal your EDA. </p>\n\n<p>I used your code to visualize the missed tracks in the volume 7 using one of our single model prediction. A question, there were some \"weird\" missed tracks, is it because those tracks were crossing -p -&gt; p and was not taken into account by visualization? \n<a href=\"https://drive.google.com/open?id=1tTEcwflL3WNvhLueb-ie0Ca7f_FVnNrX\">https://drive.google.com/open?id=1tTEcwflL3WNvhLueb-ie0Ca7f_FVnNrX</a></p>",
      "rawMarkdown": "Matthew the portion is in the volume 7, for volume 9, it should be z &lt; -500. And my approximate scoring on the volume 8 is `np.absolute(z) &lt; 500`. \n\n@Heng the portion scoring function is a bit optimistic since it doesn't consider the factor of \"losing majority\" of the extended track but merely takes the portion of detected hits from the dbscan cluster and scores on it.  Please correct me if I was wrong. However, it serves its purpose well and gave us a very good insight. Thanks for letting me steal your EDA. \n\nI used your code to visualize the missed tracks in the volume 7 using one of our single model prediction. A question, there were some \"weird\" missed tracks, is it because those tracks were crossing -p -&gt; p and was not taken into account by visualization? \nhttps://drive.google.com/open?id=1tTEcwflL3WNvhLueb-ie0Ca7f_FVnNrX",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 345134,
      "author_name": "nicolefinnie",
      "author_url": "",
      "post_date": "06/19/2018 08:47:03",
      "content": "<p>@Heng, I encountered problems when designing an LSTM so I stopped. Some coordinates have almost the same directions when their radius and angular distance are almost the same. Sometimes both of them belong to the same track (one may be a reconstructed hit), and sometimes they belong to different tracks.  Do you think an LSTM can tell them apart? I don't see how it can be more accurate, it would require heavy post processing just like the current clustering approach does. Has any of your DL models outperformed your clustering result?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 345216,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/19/2018 12:18:19",
      "content": "<p>i haven't study lstm in detail yet. this is my results on test (not train):</p>\n\n<p>left:  x,y,z\nright: r,angle,z</p>\n\n<p>gray: ground truth.</p>\n\n<p>color: lstm results</p>\n\n<p>reference: \n<a href=\"https://indico.cern.ch/event/658267/contributions/2881175/attachments/1621912/2581064/Farrell_heptrkx_ctd2018.pdf\">https://indico.cern.ch/event/658267/contributions/2881175/attachments/1621912/2581064/Farrell_heptrkx_ctd2018.pdf</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/345216/9641/LSTM_test.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 345514,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/20/2018 01:53:46",
      "content": "<p>i did some experiments. learning based lstm is better than my track extension code.</p>\n\n<p>(and i think learning-based ensemble method will work better than hand crafted ones too, and i am working on that too)</p>",
      "votes": null,
      "replies": [
        {
          "id": 346774,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/22/2018 11:00:53",
          "content": "<p>Thanks for sharing, improving clustering postprocessing is also what I am looking at now.  Looks like the next low hanging fruit. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 346809,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/22/2018 13:22:50",
      "content": "<p>using dbscan for seeding tracks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 346852,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/22/2018 15:17:37",
      "content": "<p>updated code to analyse missed and false positive seeding tracks</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/346852/9666/error2.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/346852/9667/error1.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 347312,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "06/23/2018 22:03:36",
          "content": "<p>In the script provided, you seed hits from the region z &gt; 500 and r &lt; 50, right? This captures a portion of detectors in volume 7 and 9. Why do you only seed hits from the end caps and not the 'barrel'-type detectors (use no z constraint)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347597,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "06/24/2018 19:56:46",
          "content": "<p>@Matthew the portion is in the volume 7, for volume 9, it should be z &lt; -500. And my approximate scoring on the volume 8 is <code>np.absolute(z) &lt; 500</code>. </p>\n\n<p>@Heng the portion scoring function is a bit optimistic since it doesn't consider the factor of \"losing majority\" of the extended track but merely takes the portion of detected hits from the dbscan cluster and scores on it.  Please correct me if I was wrong. However, it serves its purpose well and gave us a very good insight. Thanks for letting me steal your EDA. </p>\n\n<p>I used your code to visualize the missed tracks in the volume 7 using one of our single model prediction. A question, there were some \"weird\" missed tracks, is it because those tracks were crossing -p -&gt; p and was not taken into account by visualization? \n<a href=\"https://drive.google.com/open?id=1tTEcwflL3WNvhLueb-ie0Ca7f_FVnNrX\">https://drive.google.com/open?id=1tTEcwflL3WNvhLueb-ie0Ca7f_FVnNrX</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "345100": "has anyone start using lstm yet? lstm seems to be a potential candidate to replace kalman filter. Anyone has any results with lstm?",
    "345134": "Heng, I encountered problems when designing an LSTM so I stopped. Some coordinates have almost the same directions when their radius and angular distance are almost the same. Sometimes both of them belong to the same track (one may be a reconstructed hit), and sometimes they belong to different tracks.  Do you think an LSTM can tell them apart? I don't see how it can be more accurate, it would require heavy post processing just like the current clustering approach does. Has any of your DL models outperformed your clustering result?",
    "345216": "i haven't study lstm in detail yet. this is my results on test (not train):\n\nleft:  x,y,z\nright: r,angle,z\n\ngray: ground truth.\n\ncolor: lstm results\n\nreference: \nhttps://indico.cern.ch/event/658267/contributions/2881175/attachments/1621912/2581064/Farrell_heptrkx_ctd2018.pdf\n\n\n ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/345216/9641/LSTM_test.png",
    "345514": "i did some experiments. learning based lstm is better than my track extension code.\n\n(and i think learning-based ensemble method will work better than hand crafted ones too, and i am working on that too)",
    "346774": "Thanks for sharing, improving clustering postprocessing is also what I am looking at now.  Looks like the next low hanging fruit.",
    "346809": "using dbscan for seeding tracks",
    "346852": "updated code to analyse missed and false positive seeding tracks\n\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/346852/9666/error2.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/346852/9667/error1.png",
    "347312": "In the script provided, you seed hits from the region z &gt; 500 and r &lt; 50, right? This captures a portion of detectors in volume 7 and 9. Why do you only seed hits from the end caps and not the 'barrel'-type detectors (use no z constraint)?",
    "347597": "Matthew the portion is in the volume 7, for volume 9, it should be z &lt; -500. And my approximate scoring on the volume 8 is `np.absolute(z) &lt; 500`. \n\n@Heng the portion scoring function is a bit optimistic since it doesn't consider the factor of \"losing majority\" of the extended track but merely takes the portion of detected hits from the dbscan cluster and scores on it.  Please correct me if I was wrong. However, it serves its purpose well and gave us a very good insight. Thanks for letting me steal your EDA. \n\nI used your code to visualize the missed tracks in the volume 7 using one of our single model prediction. A question, there were some \"weird\" missed tracks, is it because those tracks were crossing -p -&gt; p and was not taken into account by visualization? \nhttps://drive.google.com/open?id=1tTEcwflL3WNvhLueb-ie0Ca7f_FVnNrX"
  },
  "source": "meta"
}