{
  "id": 58473,
  "title": "Is \"patching short track sequences\" a good idea?",
  "url": "/competitions/trackml-particle-identification/discussion/58473",
  "author_name": "Trian",
  "post_date": "2018-06-08T14:21:36.894000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Suppose you use 10x Hdbscan (or another algorithm) to get 10x 20.000 suggested tracks for an event. And suppose you get rather short tracks (mostly around 4-6 hits long, because your algorithm detects those short track sequences well). The latter is the case for my algorithm.</p>\n\n<p>Were you in a similar situation, where you THEN tried to combine, 2 tracks of length 5, by looking at overlaps (of e.g. 2 hits) and then patching the 2 tracks into 1 of length 10-2=8 hits?</p>\n\n<p>Just looking for some thoughts, before I dive deeper into this.</p>",
  "messages": [
    {
      "id": 345294,
      "postDate": "2018-06-19T15:58:01.223Z",
      "content": "<p>I've been trying to combine tracks and all I've accomplished is making the score worse.  I've tried a couple things, but the best (of bad) results came by fitting a cylinder to the track points and getting the direction vector (axis of the cylinder), the radius of the cylinder (which is problematic with \"straight\" lines and planar curves), and the endpoints.  I used nearest neighbors to cluster the tracks using the direction, radius and first and last points (i.e., I fit the model using direction, radius and z-coord of the last point,  then I matched nearest neighbor using direction, radius and z-coord of the first point).  I may come back to \"combining\" tracks later, but like <a href=\"/macfarll\">@macfarll</a> pointed out, the track extension code works much better for now.</p>",
      "rawMarkdown": "I've been trying to combine tracks and all I've accomplished is making the score worse.  I've tried a couple things, but the best (of bad) results came by fitting a cylinder to the track points and getting the direction vector (axis of the cylinder), the radius of the cylinder (which is problematic with \"straight\" lines and planar curves), and the endpoints.  I used nearest neighbors to cluster the tracks using the direction, radius and first and last points (i.e., I fit the model using direction, radius and z-coord of the last point,  then I matched nearest neighbor using direction, radius and z-coord of the first point).  I may come back to \"combining\" tracks later, but like @macfarll pointed out, the track extension code works much better for now.",
      "votes": 1,
      "replies": [
        {
          "id": 347144,
          "postDate": "2018-06-23T11:21:01.170Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 340172,
      "postDate": "2018-06-08T14:21:36.893Z",
      "content": "<p>Suppose you use 10x Hdbscan (or another algorithm) to get 10x 20.000 suggested tracks for an event. And suppose you get rather short tracks (mostly around 4-6 hits long, because your algorithm detects those short track sequences well). The latter is the case for my algorithm.</p>\n\n<p>Were you in a similar situation, where you THEN tried to combine, 2 tracks of length 5, by looking at overlaps (of e.g. 2 hits) and then patching the 2 tracks into 1 of length 10-2=8 hits?</p>\n\n<p>Just looking for some thoughts, before I dive deeper into this.</p>",
      "rawMarkdown": "Suppose you use 10x Hdbscan (or another algorithm) to get 10x 20.000 suggested tracks for an event. And suppose you get rather short tracks (mostly around 4-6 hits long, because your algorithm detects those short track sequences well). The latter is the case for my algorithm.\n\nWere you in a similar situation, where you THEN tried to combine, 2 tracks of length 5, by looking at overlaps (of e.g. 2 hits) and then patching the 2 tracks into 1 of length 10-2=8 hits?\n\nJust looking for some thoughts, before I dive deeper into this.",
      "votes": 1
    },
    {
      "id": 345287,
      "postDate": "2018-06-19T15:29:17.320Z",
      "content": "<p>I'd recommend you look into track extension instead, since most recognized tracks probably won't overlap. The extension code will do similar to what you described, but is more likely to get the endpoints that the clustering missed, which appears to be an issue with clustering (I believe this is because high values of z are the most impacted by the helix shape, low values are more sensitive to noise).</p>\n\n<p>From what I've seen, if anything dbscan will be likely to merge two tracks that don't belong together than it is to label two parts of a track independently. </p>\n\n<p>It is worth noting I haven't really been looking for tracks that are split...</p>",
      "rawMarkdown": "I'd recommend you look into track extension instead, since most recognized tracks probably won't overlap. The extension code will do similar to what you described, but is more likely to get the endpoints that the clustering missed, which appears to be an issue with clustering (I believe this is because high values of z are the most impacted by the helix shape, low values are more sensitive to noise).\n\nFrom what I've seen, if anything dbscan will be likely to merge two tracks that don't belong together than it is to label two parts of a track independently. \n\nIt is worth noting I haven't really been looking for tracks that are split...",
      "replies": [
        {
          "id": 347145,
          "postDate": "2018-06-23T11:21:21.953Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 345294,
      "author_name": "Michael Maguire",
      "author_url": "",
      "post_date": "2018-06-19T15:58:01.223000",
      "content": "<p>I've been trying to combine tracks and all I've accomplished is making the score worse.  I've tried a couple things, but the best (of bad) results came by fitting a cylinder to the track points and getting the direction vector (axis of the cylinder), the radius of the cylinder (which is problematic with \"straight\" lines and planar curves), and the endpoints.  I used nearest neighbors to cluster the tracks using the direction, radius and first and last points (i.e., I fit the model using direction, radius and z-coord of the last point,  then I matched nearest neighbor using direction, radius and z-coord of the first point).  I may come back to \"combining\" tracks later, but like <a href=\"/macfarll\">@macfarll</a> pointed out, the track extension code works much better for now.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 347144,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-23T11:21:01.170000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 345287,
      "author_name": "macfarll",
      "author_url": "",
      "post_date": "2018-06-19T15:29:17.320000",
      "content": "<p>I'd recommend you look into track extension instead, since most recognized tracks probably won't overlap. The extension code will do similar to what you described, but is more likely to get the endpoints that the clustering missed, which appears to be an issue with clustering (I believe this is because high values of z are the most impacted by the helix shape, low values are more sensitive to noise).</p>\n\n<p>From what I've seen, if anything dbscan will be likely to merge two tracks that don't belong together than it is to label two parts of a track independently. </p>\n\n<p>It is worth noting I haven't really been looking for tracks that are split...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 347145,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-23T11:21:21.953000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "345294": "I've been trying to combine tracks and all I've accomplished is making the score worse.  I've tried a couple things, but the best (of bad) results came by fitting a cylinder to the track points and getting the direction vector (axis of the cylinder), the radius of the cylinder (which is problematic with \"straight\" lines and planar curves), and the endpoints.  I used nearest neighbors to cluster the tracks using the direction, radius and first and last points (i.e., I fit the model using direction, radius and z-coord of the last point,  then I matched nearest neighbor using direction, radius and z-coord of the first point).  I may come back to \"combining\" tracks later, but like @macfarll pointed out, the track extension code works much better for now.",
    "340172": "Suppose you use 10x Hdbscan (or another algorithm) to get 10x 20.000 suggested tracks for an event. And suppose you get rather short tracks (mostly around 4-6 hits long, because your algorithm detects those short track sequences well). The latter is the case for my algorithm.\n\nWere you in a similar situation, where you THEN tried to combine, 2 tracks of length 5, by looking at overlaps (of e.g. 2 hits) and then patching the 2 tracks into 1 of length 10-2=8 hits?\n\nJust looking for some thoughts, before I dive deeper into this.",
    "345287": "I'd recommend you look into track extension instead, since most recognized tracks probably won't overlap. The extension code will do similar to what you described, but is more likely to get the endpoints that the clustering missed, which appears to be an issue with clustering (I believe this is because high values of z are the most impacted by the helix shape, low values are more sensitive to noise).\n\nFrom what I've seen, if anything dbscan will be likely to merge two tracks that don't belong together than it is to label two parts of a track independently. \n\nIt is worth noting I haven't really been looking for tracks that are split..."
  }
}