{
  "id": 60680,
  "title": "DBSCAN using pair as input",
  "url": "/competitions/trackml-particle-identification/discussion/60680",
  "author_name": "",
  "post_date": "2018-07-08T10:23:49.610008600Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Here is the code and results.</p>\n\n<p>Method:</p>\n\n<ol>\n<li><p>make all possible pairs for hits on two consecutive layers.</p></li>\n<li><p>for each pair, compute the unit normal vector.</p></li>\n<li><p>the pair is represented as: pair = { start_hit, end_hit, normal vector}</p></li>\n<li><p>compute pair distance as dist( pair_a, pair b) = INF if pair_a.end_hit != pair_b.start_hit, else\ndist( pair_a, pair b) = || pair_a.normal - pair_b.normal ||</p></li>\n<li><p>use DBSCAN clustering  using precomputed distances</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/353941/9816/Figure_3-5.png\" alt=\"enter image description here\"></p></li>\n</ol>\n\n<p>BLACK: ground truth</p>\n\n<p>COLOR : DBSCAN on pairs</p>",
  "messages": [
    {
      "id": "353941",
      "postDate": "07/08/2018 10:23:49",
      "content": "<p>Here is the code and results.</p>\n\n<p>Method:</p>\n\n<ol>\n<li><p>make all possible pairs for hits on two consecutive layers.</p></li>\n<li><p>for each pair, compute the unit normal vector.</p></li>\n<li><p>the pair is represented as: pair = { start_hit, end_hit, normal vector}</p></li>\n<li><p>compute pair distance as dist( pair_a, pair b) = INF if pair_a.end_hit != pair_b.start_hit, else\ndist( pair_a, pair b) = || pair_a.normal - pair_b.normal ||</p></li>\n<li><p>use DBSCAN clustering  using precomputed distances</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/353941/9816/Figure_3-5.png\" alt=\"enter image description here\"></p></li>\n</ol>\n\n<p>BLACK: ground truth</p>\n\n<p>COLOR : DBSCAN on pairs</p>",
      "rawMarkdown": "Here is the code and results.\n\nMethod:\n\n1. make all possible pairs for hits on two consecutive layers.\n\n2. for each pair, compute the unit normal vector.\n\n3. the pair is represented as: pair = { start_hit, end_hit, normal vector}\n\n4. compute pair distance as dist( pair_a, pair b) = INF if pair_a.end_hit != pair_b.start_hit, else\ndist( pair_a, pair b) = || pair_a.normal - pair_b.normal ||\n\n5. use DBSCAN clustering  using precomputed distances\n\n  ![enter image description here][1]\n\n\nBLACK: ground truth\n\nCOLOR : DBSCAN on pairs\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/353941/9816/Figure_3-5.png",
      "votes": null
    },
    {
      "id": "353976",
      "postDate": "07/08/2018 12:07:13",
      "content": "<p>Interesting thought, as always from you.  Just one remark, the vector you compute is a direction vector for the line going through the pair, not a normal vector.  A normal vector is perpendicular to the line going through the pair.  </p>",
      "rawMarkdown": "Interesting thought, as always from you.  Just one remark, the vector you compute is a direction vector for the line going through the pair, not a normal vector.  A normal vector is perpendicular to the line going through the pair.",
      "votes": null
    },
    {
      "id": "353980",
      "postDate": "07/08/2018 12:15:16",
      "content": "<p>google for \"tracking cellular automaton\"</p>",
      "rawMarkdown": "google for \"tracking cellular automaton\"",
      "votes": null
    },
    {
      "id": "354036",
      "postDate": "07/08/2018 14:54:03",
      "content": "<p>@cpmp Thanks. You are right. It should be unit direction or tangent vector</p>",
      "rawMarkdown": "cpmp Thanks. You are right. It should be unit direction or tangent vector",
      "votes": null
    },
    {
      "id": "356846",
      "postDate": "07/14/2018 15:44:22",
      "content": "<p>I did something similar but didn't finish it (got too busy so I stopped). I was hoping to build an approach without unrolling</p>\n\n<p>1) I selected pairs but without limiting hits to consecutives layers as you did. I used a number of event files to learn most likely distribution of links between volume/layer/modules then used it to select candidate neighboors for each hit from a test event and ended up with 26mil pairs (still a lot I believe but much less than 120k**2).</p>\n\n<p>2) Then I assumed each pair as a triplet with the origin (0,0) from the x/y plane</p>\n\n<p>3) From this, I generated features for each pair and used the helix equation to extract its parameters. Technically, if the tracks were near-perfect helixes, I should get similar parameters, even considering (x0,y0)=(0,0) as the origin while z0 can vary. This was combined with other single-hit features with lesser weight (such as z2=z/r_xy). The origin z0 was also estimated (this one is trickly, because a helix can have multiple acceptable solutions given certain hits)</p>\n\n<p>4) DBSCAN would take too long, so I started a custom approach to build tracks iteratively as I go thru the pair-wise hits file.</p>\n\n<p>Unfortunately, I found way too many pairs with very similar features which ended up clustering together. It seems like a big part of the problem is related to the fact that most tracks are poor approximations of true helixes/arcs (using truth file, I could confirm this). So I had to use a larger variance/epsilon for the pair-wise distance of the pair-wise hits, which explains the poor results. Or my helix approximation was wrong (but I tried it on simulated helix data and it reproduces the parameters perfectly)...</p>",
      "rawMarkdown": "I did something similar but didn't finish it (got too busy so I stopped). I was hoping to build an approach without unrolling\n\n1) I selected pairs but without limiting hits to consecutives layers as you did. I used a number of event files to learn most likely distribution of links between volume/layer/modules then used it to select candidate neighboors for each hit from a test event and ended up with 26mil pairs (still a lot I believe but much less than 120k**2).\n\n2) Then I assumed each pair as a triplet with the origin (0,0) from the x/y plane\n\n3) From this, I generated features for each pair and used the helix equation to extract its parameters. Technically, if the tracks were near-perfect helixes, I should get similar parameters, even considering (x0,y0)=(0,0) as the origin while z0 can vary. This was combined with other single-hit features with lesser weight (such as z2=z/r_xy). The origin z0 was also estimated (this one is trickly, because a helix can have multiple acceptable solutions given certain hits)\n\n4) DBSCAN would take too long, so I started a custom approach to build tracks iteratively as I go thru the pair-wise hits file.\n\nUnfortunately, I found way too many pairs with very similar features which ended up clustering together. It seems like a big part of the problem is related to the fact that most tracks are poor approximations of true helixes/arcs (using truth file, I could confirm this). So I had to use a larger variance/epsilon for the pair-wise distance of the pair-wise hits, which explains the poor results. Or my helix approximation was wrong (but I tried it on simulated helix data and it reproduces the parameters perfectly)...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 353976,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/08/2018 12:07:13",
      "content": "<p>Interesting thought, as always from you.  Just one remark, the vector you compute is a direction vector for the line going through the pair, not a normal vector.  A normal vector is perpendicular to the line going through the pair.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 353980,
      "author_name": "vininn",
      "author_url": "",
      "post_date": "07/08/2018 12:15:16",
      "content": "<p>google for \"tracking cellular automaton\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 354036,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/08/2018 14:54:03",
      "content": "<p>@cpmp Thanks. You are right. It should be unit direction or tangent vector</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 356846,
      "author_name": "riadsouissi",
      "author_url": "",
      "post_date": "07/14/2018 15:44:22",
      "content": "<p>I did something similar but didn't finish it (got too busy so I stopped). I was hoping to build an approach without unrolling</p>\n\n<p>1) I selected pairs but without limiting hits to consecutives layers as you did. I used a number of event files to learn most likely distribution of links between volume/layer/modules then used it to select candidate neighboors for each hit from a test event and ended up with 26mil pairs (still a lot I believe but much less than 120k**2).</p>\n\n<p>2) Then I assumed each pair as a triplet with the origin (0,0) from the x/y plane</p>\n\n<p>3) From this, I generated features for each pair and used the helix equation to extract its parameters. Technically, if the tracks were near-perfect helixes, I should get similar parameters, even considering (x0,y0)=(0,0) as the origin while z0 can vary. This was combined with other single-hit features with lesser weight (such as z2=z/r_xy). The origin z0 was also estimated (this one is trickly, because a helix can have multiple acceptable solutions given certain hits)</p>\n\n<p>4) DBSCAN would take too long, so I started a custom approach to build tracks iteratively as I go thru the pair-wise hits file.</p>\n\n<p>Unfortunately, I found way too many pairs with very similar features which ended up clustering together. It seems like a big part of the problem is related to the fact that most tracks are poor approximations of true helixes/arcs (using truth file, I could confirm this). So I had to use a larger variance/epsilon for the pair-wise distance of the pair-wise hits, which explains the poor results. Or my helix approximation was wrong (but I tried it on simulated helix data and it reproduces the parameters perfectly)...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "353941": "Here is the code and results.\n\nMethod:\n\n1. make all possible pairs for hits on two consecutive layers.\n\n2. for each pair, compute the unit normal vector.\n\n3. the pair is represented as: pair = { start_hit, end_hit, normal vector}\n\n4. compute pair distance as dist( pair_a, pair b) = INF if pair_a.end_hit != pair_b.start_hit, else\ndist( pair_a, pair b) = || pair_a.normal - pair_b.normal ||\n\n5. use DBSCAN clustering  using precomputed distances\n\n  ![enter image description here][1]\n\n\nBLACK: ground truth\n\nCOLOR : DBSCAN on pairs\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/353941/9816/Figure_3-5.png",
    "353976": "Interesting thought, as always from you.  Just one remark, the vector you compute is a direction vector for the line going through the pair, not a normal vector.  A normal vector is perpendicular to the line going through the pair.",
    "353980": "google for \"tracking cellular automaton\"",
    "354036": "cpmp Thanks. You are right. It should be unit direction or tangent vector",
    "356846": "I did something similar but didn't finish it (got too busy so I stopped). I was hoping to build an approach without unrolling\n\n1) I selected pairs but without limiting hits to consecutives layers as you did. I used a number of event files to learn most likely distribution of links between volume/layer/modules then used it to select candidate neighboors for each hit from a test event and ended up with 26mil pairs (still a lot I believe but much less than 120k**2).\n\n2) Then I assumed each pair as a triplet with the origin (0,0) from the x/y plane\n\n3) From this, I generated features for each pair and used the helix equation to extract its parameters. Technically, if the tracks were near-perfect helixes, I should get similar parameters, even considering (x0,y0)=(0,0) as the origin while z0 can vary. This was combined with other single-hit features with lesser weight (such as z2=z/r_xy). The origin z0 was also estimated (this one is trickly, because a helix can have multiple acceptable solutions given certain hits)\n\n4) DBSCAN would take too long, so I started a custom approach to build tracks iteratively as I go thru the pair-wise hits file.\n\nUnfortunately, I found way too many pairs with very similar features which ended up clustering together. It seems like a big part of the problem is related to the fact that most tracks are poor approximations of true helixes/arcs (using truth file, I could confirm this). So I had to use a larger variance/epsilon for the pair-wise distance of the pair-wise hits, which explains the poor results. Or my helix approximation was wrong (but I tried it on simulated helix data and it reproduces the parameters perfectly)..."
  },
  "source": "meta"
}