{
  "id": 57704,
  "title": "DBSCAN doesn't group together hits from the same particle",
  "url": "/competitions/trackml-particle-identification/discussion/57704",
  "author_name": "",
  "post_date": "2018-05-28T00:41:12.418051800Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I re-implemented the benchmark DBSCAN kernel <a href=\"https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\">here</a> and got a score of around 0.2. After viewing the ground truth tracks and the predicted tracks I found that the DBSCAN algorithm doesn't group together hits from the same particle very well.  For example: </p>\n\n<p>The ground truth tracks look like this:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/334597/9532/ground-truth-tracks.png\" alt=\"Ground truth tracks\"></p>\n\n<p>And the predicted tracks look like this:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/334597/9533/predicted-tracks.png\" alt=\"\"></p>\n\n<p>Is this a implementation mistake of my behalf,  or is it a limitation of the DBSCAN algorithm, or do the hits need to be pre-processed using a different method that would group hits from a particle close together and hits from different particles farther away? </p>\n\n<p>Thanks</p>\n\n<hr>\n\n<h3>Please Note:</h3>\n\n<ol>\n<li>The predicted tracks image only shows tracks that have a minimum of two hits on them</li>\n<li>Although the tracks in the predicted tracks image look like points, they are each lines between a minimum of two hits</li>\n<li>The tracks have been transformed and scaled by the same equations which are in the benchmark kernel</li>\n</ol>",
  "messages": [
    {
      "id": "334597",
      "postDate": "05/28/2018 00:41:12",
      "content": "<p>Hi,</p>\n\n<p>I re-implemented the benchmark DBSCAN kernel <a href=\"https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\">here</a> and got a score of around 0.2. After viewing the ground truth tracks and the predicted tracks I found that the DBSCAN algorithm doesn't group together hits from the same particle very well.  For example: </p>\n\n<p>The ground truth tracks look like this:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/334597/9532/ground-truth-tracks.png\" alt=\"Ground truth tracks\"></p>\n\n<p>And the predicted tracks look like this:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/334597/9533/predicted-tracks.png\" alt=\"\"></p>\n\n<p>Is this a implementation mistake of my behalf,  or is it a limitation of the DBSCAN algorithm, or do the hits need to be pre-processed using a different method that would group hits from a particle close together and hits from different particles farther away? </p>\n\n<p>Thanks</p>\n\n<hr>\n\n<h3>Please Note:</h3>\n\n<ol>\n<li>The predicted tracks image only shows tracks that have a minimum of two hits on them</li>\n<li>Although the tracks in the predicted tracks image look like points, they are each lines between a minimum of two hits</li>\n<li>The tracks have been transformed and scaled by the same equations which are in the benchmark kernel</li>\n</ol>",
      "rawMarkdown": "Hi,\n\nI re-implemented the benchmark DBSCAN kernel [here][1] and got a score of around 0.2. After viewing the ground truth tracks and the predicted tracks I found that the DBSCAN algorithm doesn't group together hits from the same particle very well.  For example: \n\nThe ground truth tracks look like this:\n\n![Ground truth tracks][2]\n\nAnd the predicted tracks look like this:\n\n![][3]\n\nIs this a implementation mistake of my behalf,  or is it a limitation of the DBSCAN algorithm, or do the hits need to be pre-processed using a different method that would group hits from a particle close together and hits from different particles farther away? \n\nThanks\n ---\n### Please Note: \n\n1. The predicted tracks image only shows tracks that have a minimum of two hits on them\n2. Although the tracks in the predicted tracks image look like points, they are each lines between a minimum of two hits\n3. The tracks have been transformed and scaled by the same equations which are in the benchmark kernel\n\n[1]: https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\n[2]: https://storage.googleapis.com/kaggle-forum-message-attachments/334597/9532/ground-truth-tracks.png\n[3]: https://storage.googleapis.com/kaggle-forum-message-attachments/334597/9533/predicted-tracks.png",
      "votes": null
    },
    {
      "id": "334656",
      "postDate": "05/28/2018 05:31:07",
      "content": "<p>It looks like you're plotting both in transformed coordinates. While you are probably clustering with the transformed coordinates, but your model does a better job at predicting the straight lines. The predicted tracks appear as a single point, because they're so close in transformed space. The tracks that rotate around the z axis in the truth tracks are getting missed, but those tracks are not the ones your model is representing. </p>",
      "rawMarkdown": "It looks like you're plotting both in transformed coordinates. While you are probably clustering with the transformed coordinates, but your model does a better job at predicting the straight lines. The predicted tracks appear as a single point, because they're so close in transformed space. The tracks that rotate around the z axis in the truth tracks are getting missed, but those tracks are not the ones your model is representing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 334656,
      "author_name": "macfarll",
      "author_url": "",
      "post_date": "05/28/2018 05:31:07",
      "content": "<p>It looks like you're plotting both in transformed coordinates. While you are probably clustering with the transformed coordinates, but your model does a better job at predicting the straight lines. The predicted tracks appear as a single point, because they're so close in transformed space. The tracks that rotate around the z axis in the truth tracks are getting missed, but those tracks are not the ones your model is representing. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "334597": "Hi,\n\nI re-implemented the benchmark DBSCAN kernel [here][1] and got a score of around 0.2. After viewing the ground truth tracks and the predicted tracks I found that the DBSCAN algorithm doesn't group together hits from the same particle very well.  For example: \n\nThe ground truth tracks look like this:\n\n![Ground truth tracks][2]\n\nAnd the predicted tracks look like this:\n\n![][3]\n\nIs this a implementation mistake of my behalf,  or is it a limitation of the DBSCAN algorithm, or do the hits need to be pre-processed using a different method that would group hits from a particle close together and hits from different particles farther away? \n\nThanks\n ---\n### Please Note: \n\n1. The predicted tracks image only shows tracks that have a minimum of two hits on them\n2. Although the tracks in the predicted tracks image look like points, they are each lines between a minimum of two hits\n3. The tracks have been transformed and scaled by the same equations which are in the benchmark kernel\n\n[1]: https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\n[2]: https://storage.googleapis.com/kaggle-forum-message-attachments/334597/9532/ground-truth-tracks.png\n[3]: https://storage.googleapis.com/kaggle-forum-message-attachments/334597/9533/predicted-tracks.png",
    "334656": "It looks like you're plotting both in transformed coordinates. While you are probably clustering with the transformed coordinates, but your model does a better job at predicting the straight lines. The predicted tracks appear as a single point, because they're so close in transformed space. The tracks that rotate around the z axis in the truth tracks are getting missed, but those tracks are not the ones your model is representing."
  },
  "source": "meta"
}