{
  "id": 56040,
  "title": "Instructions to calculate track_id",
  "url": "/competitions/trackml-particle-identification/discussion/56040",
  "author_name": "",
  "post_date": "2018-05-04T21:59:02.276031100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I apologize if this is answered somewhere, but I have looked and just don't see it:</p>\n\n<p>How do I calculate the track_id?  I am using R, so I do not have access to the Python package.  I am happy to code up something custom, and have already got the train and test set built out.  But I am missing the actual label I need to train against, which I assume is calculated from the supplied data.</p>\n\n<p>Or am I missing something?</p>",
  "messages": [
    {
      "id": "323346",
      "postDate": "05/04/2018 21:59:02",
      "content": "<p>I apologize if this is answered somewhere, but I have looked and just don't see it:</p>\n\n<p>How do I calculate the track_id?  I am using R, so I do not have access to the Python package.  I am happy to code up something custom, and have already got the train and test set built out.  But I am missing the actual label I need to train against, which I assume is calculated from the supplied data.</p>\n\n<p>Or am I missing something?</p>",
      "rawMarkdown": "I apologize if this is answered somewhere, but I have looked and just don't see it:\n\nHow do I calculate the track_id?  I am using R, so I do not have access to the Python package.  I am happy to code up something custom, and have already got the train and test set built out.  But I am missing the actual label I need to train against, which I assume is calculated from the supplied data.\n\nOr am I missing something?",
      "votes": null
    },
    {
      "id": "323348",
      "postDate": "05/04/2018 22:11:12",
      "content": "<p>Here is what I have gathered so far, but may be confused:</p>\n\n<p>we have points of origin and points of termination to train on.  We could build a model to predict points of termination from points of origin.  </p>\n\n<p>Next, we could imagine there are pieces of yarn connecting each origin and termination point, where the thickness of the yarn accounts for multiple paths.  It's basically discretization of the space, where every path within some radius is counted as the same path.  This will group multiple data points into the same yarn path.</p>\n\n<p>Let's then name each piece of yarn with a distinct integer moniker.  It doesn't matter the logic we follow.  We could start with 0 and count by 1, or count by primes -- all that matters is that we are consistent.</p>\n\n<p>This yarn # is what we submit.  The judges of the competition will then be able to tell whether we grouped the correct particles into the same collection, thereby showing we matched up their path.</p>\n\n<p>I am not sure why this is better than just measuring the distance from what was predicted and what was true, but this is what I have pieced together ....</p>",
      "rawMarkdown": "Here is what I have gathered so far, but may be confused:\n\nwe have points of origin and points of termination to train on.  We could build a model to predict points of termination from points of origin.  \n\nNext, we could imagine there are pieces of yarn connecting each origin and termination point, where the thickness of the yarn accounts for multiple paths.  It's basically discretization of the space, where every path within some radius is counted as the same path.  This will group multiple data points into the same yarn path.\n\nLet's then name each piece of yarn with a distinct integer moniker.  It doesn't matter the logic we follow.  We could start with 0 and count by 1, or count by primes -- all that matters is that we are consistent.\n\nThis yarn # is what we submit.  The judges of the competition will then be able to tell whether we grouped the correct particles into the same collection, thereby showing we matched up their path.\n\nI am not sure why this is better than just measuring the distance from what was predicted and what was true, but this is what I have pieced together ....",
      "votes": null
    },
    {
      "id": "323621",
      "postDate": "05/05/2018 17:25:27",
      "content": "<p>In the solution, the track_id are completely arbitrary.  You number the tracks as you wish, provided that:</p>\n\n<ul>\n<li>The reconstructed tracks are uniquely defined within each\nevent. </li>\n<li><p>They belong to [0, 1E9]</p>\n\n<p>The scoring function takes care of associating your tracks with the true ones from the hits they contain. For more details, please look at the evaluation tab (in overview) or the score section <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">here</a> </p></li>\n</ul>",
      "rawMarkdown": "In the solution, the track_id are completely arbitrary.  You number the tracks as you wish, provided that:\n \n\n - The reconstructed tracks are uniquely defined within each\n   event. \n - They belong to [0, 1E9]\n\n The scoring function takes care of associating your tracks with the true ones from the hits they contain. For more details, please look at the evaluation tab (in overview) or the score section [here][1] \n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf",
      "votes": null
    },
    {
      "id": "323647",
      "postDate": "05/05/2018 18:56:44",
      "content": "<p>Thanks for being willing to help.  I am sure your explanation is correct and well stated, but I need an explanation that takes a step back and does not require me to already understand the problem.  This is not my field.  I did read the instructions but find it presumes understanding that comes from being in the field.</p>\n\n<p>We are given x,y,z of the origin and x,y,z of different stages, including the endpoint.  I could predict the x,y,z of each stagegate until the end.  That would be a 'track', and every particle would have it's own unique 'track', excepting coincidence.  Every particle would have it's own track_id, right?</p>\n\n<p>If I were to create a fuzzy boundary around the tracks I could cluster them and use unique ids of my own creation.   I don't understand how this is helpful.</p>\n\n<p>Is this the idea - a two step problem that starts with a series of regressions, followed by clustering?</p>",
      "rawMarkdown": "Thanks for being willing to help.  I am sure your explanation is correct and well stated, but I need an explanation that takes a step back and does not require me to already understand the problem.  This is not my field.  I did read the instructions but find it presumes understanding that comes from being in the field.\n\nWe are given x,y,z of the origin and x,y,z of different stages, including the endpoint.  I could predict the x,y,z of each stagegate until the end.  That would be a 'track', and every particle would have it's own unique 'track', excepting coincidence.  Every particle would have it's own track_id, right?\n\nIf I were to create a fuzzy boundary around the tracks I could cluster them and use unique ids of my own creation.   I don't understand how this is helpful.\n\nIs this the idea - a two step problem that starts with a series of regressions, followed by clustering?",
      "votes": null
    },
    {
      "id": "326241",
      "postDate": "05/09/2018 13:26:46",
      "content": "<p>@SonicAlch3mist - we are NOT given the point of origin. </p>\n\n<p>It might be confusing because the training files contain much more information than the information provided for testing.  </p>\n\n<p>For example \"<em>-particles.csv\" files are not provided in the test set. It's true that you have the \"origin\" in that files but you don't have that information at inference time. At inference you only have the data from \"</em>-hits.csv\" and \"*-cells.csv\" files. </p>\n\n<p>So, we are given various points (hits) that represent locations where particles where detected. </p>\n\n<p>We need to find sub-sets of points that were generated by each particle. </p>\n\n<p>We can call this sub-sets of points as \"tracks\", because if you join the points in a sub set it would represent how the \"particle\" moved through the detectors. </p>",
      "rawMarkdown": "SonicAlch3mist - we are NOT given the point of origin. \n\nIt might be confusing because the training files contain much more information than the information provided for testing.  \n\nFor example \"*-particles.csv\" files are not provided in the test set. It's true that you have the \"origin\" in that files but you don't have that information at inference time. At inference you only have the data from \"*-hits.csv\" and \"*-cells.csv\" files. \n\nSo, we are given various points (hits) that represent locations where particles where detected. \n\nWe need to find sub-sets of points that were generated by each particle. \n\nWe can call this sub-sets of points as \"tracks\", because if you join the points in a sub set it would represent how the \"particle\" moved through the detectors.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 323348,
      "author_name": "sonicalch3mist",
      "author_url": "",
      "post_date": "05/04/2018 22:11:12",
      "content": "<p>Here is what I have gathered so far, but may be confused:</p>\n\n<p>we have points of origin and points of termination to train on.  We could build a model to predict points of termination from points of origin.  </p>\n\n<p>Next, we could imagine there are pieces of yarn connecting each origin and termination point, where the thickness of the yarn accounts for multiple paths.  It's basically discretization of the space, where every path within some radius is counted as the same path.  This will group multiple data points into the same yarn path.</p>\n\n<p>Let's then name each piece of yarn with a distinct integer moniker.  It doesn't matter the logic we follow.  We could start with 0 and count by 1, or count by primes -- all that matters is that we are consistent.</p>\n\n<p>This yarn # is what we submit.  The judges of the competition will then be able to tell whether we grouped the correct particles into the same collection, thereby showing we matched up their path.</p>\n\n<p>I am not sure why this is better than just measuring the distance from what was predicted and what was true, but this is what I have pieced together ....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 323621,
      "author_name": "cecilegermain",
      "author_url": "",
      "post_date": "05/05/2018 17:25:27",
      "content": "<p>In the solution, the track_id are completely arbitrary.  You number the tracks as you wish, provided that:</p>\n\n<ul>\n<li>The reconstructed tracks are uniquely defined within each\nevent. </li>\n<li><p>They belong to [0, 1E9]</p>\n\n<p>The scoring function takes care of associating your tracks with the true ones from the hits they contain. For more details, please look at the evaluation tab (in overview) or the score section <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">here</a> </p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 323647,
          "author_name": "sonicalch3mist",
          "author_url": "",
          "post_date": "05/05/2018 18:56:44",
          "content": "<p>Thanks for being willing to help.  I am sure your explanation is correct and well stated, but I need an explanation that takes a step back and does not require me to already understand the problem.  This is not my field.  I did read the instructions but find it presumes understanding that comes from being in the field.</p>\n\n<p>We are given x,y,z of the origin and x,y,z of different stages, including the endpoint.  I could predict the x,y,z of each stagegate until the end.  That would be a 'track', and every particle would have it's own unique 'track', excepting coincidence.  Every particle would have it's own track_id, right?</p>\n\n<p>If I were to create a fuzzy boundary around the tracks I could cluster them and use unique ids of my own creation.   I don't understand how this is helpful.</p>\n\n<p>Is this the idea - a two step problem that starts with a series of regressions, followed by clustering?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326241,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/09/2018 13:26:46",
          "content": "<p>@SonicAlch3mist - we are NOT given the point of origin. </p>\n\n<p>It might be confusing because the training files contain much more information than the information provided for testing.  </p>\n\n<p>For example \"<em>-particles.csv\" files are not provided in the test set. It's true that you have the \"origin\" in that files but you don't have that information at inference time. At inference you only have the data from \"</em>-hits.csv\" and \"*-cells.csv\" files. </p>\n\n<p>So, we are given various points (hits) that represent locations where particles where detected. </p>\n\n<p>We need to find sub-sets of points that were generated by each particle. </p>\n\n<p>We can call this sub-sets of points as \"tracks\", because if you join the points in a sub set it would represent how the \"particle\" moved through the detectors. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "323346": "I apologize if this is answered somewhere, but I have looked and just don't see it:\n\nHow do I calculate the track_id?  I am using R, so I do not have access to the Python package.  I am happy to code up something custom, and have already got the train and test set built out.  But I am missing the actual label I need to train against, which I assume is calculated from the supplied data.\n\nOr am I missing something?",
    "323348": "Here is what I have gathered so far, but may be confused:\n\nwe have points of origin and points of termination to train on.  We could build a model to predict points of termination from points of origin.  \n\nNext, we could imagine there are pieces of yarn connecting each origin and termination point, where the thickness of the yarn accounts for multiple paths.  It's basically discretization of the space, where every path within some radius is counted as the same path.  This will group multiple data points into the same yarn path.\n\nLet's then name each piece of yarn with a distinct integer moniker.  It doesn't matter the logic we follow.  We could start with 0 and count by 1, or count by primes -- all that matters is that we are consistent.\n\nThis yarn # is what we submit.  The judges of the competition will then be able to tell whether we grouped the correct particles into the same collection, thereby showing we matched up their path.\n\nI am not sure why this is better than just measuring the distance from what was predicted and what was true, but this is what I have pieced together ....",
    "323621": "In the solution, the track_id are completely arbitrary.  You number the tracks as you wish, provided that:\n \n\n - The reconstructed tracks are uniquely defined within each\n   event. \n - They belong to [0, 1E9]\n\n The scoring function takes care of associating your tracks with the true ones from the hits they contain. For more details, please look at the evaluation tab (in overview) or the score section [here][1] \n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf",
    "323647": "Thanks for being willing to help.  I am sure your explanation is correct and well stated, but I need an explanation that takes a step back and does not require me to already understand the problem.  This is not my field.  I did read the instructions but find it presumes understanding that comes from being in the field.\n\nWe are given x,y,z of the origin and x,y,z of different stages, including the endpoint.  I could predict the x,y,z of each stagegate until the end.  That would be a 'track', and every particle would have it's own unique 'track', excepting coincidence.  Every particle would have it's own track_id, right?\n\nIf I were to create a fuzzy boundary around the tracks I could cluster them and use unique ids of my own creation.   I don't understand how this is helpful.\n\nIs this the idea - a two step problem that starts with a series of regressions, followed by clustering?",
    "326241": "SonicAlch3mist - we are NOT given the point of origin. \n\nIt might be confusing because the training files contain much more information than the information provided for testing.  \n\nFor example \"*-particles.csv\" files are not provided in the test set. It's true that you have the \"origin\" in that files but you don't have that information at inference time. At inference you only have the data from \"*-hits.csv\" and \"*-cells.csv\" files. \n\nSo, we are given various points (hits) that represent locations where particles where detected. \n\nWe need to find sub-sets of points that were generated by each particle. \n\nWe can call this sub-sets of points as \"tracks\", because if you join the points in a sub set it would represent how the \"particle\" moved through the detectors."
  },
  "source": "meta"
}