{
  "id": 62249,
  "title": "A fundamental doubt",
  "url": "/competitions/trackml-particle-identification/discussion/62249",
  "author_name": "",
  "post_date": "2018-07-30T11:58:25.715152700Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, \nSorry for a Novice question about the competition. Even after going through the official manual, some good EDA and other kernels I fail to understand the core theme of this competition. I have a few questions, could someone please help me out by explaining in simple terms.</p>\n\n<p>1) Are the tracks (that we try to predict) predefined?\n<a href=\"https://emoyse.web.cern.ch/emoyse/WebEventDisplay/jsdisplay_TrackML.html\">The 3D viewer</a> provided me a better visualisation of the problem, but I don't get a clear picture of what we are trying to predict.</p>\n\n<p>2) What exactly are the 'labels' that we are predicting using the standard DBScan method mentioned in the <a href=\"https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\">official kernels</a> ?</p>",
  "messages": [
    {
      "id": "363951",
      "postDate": "07/30/2018 11:58:25",
      "content": "<p>Hi, \nSorry for a Novice question about the competition. Even after going through the official manual, some good EDA and other kernels I fail to understand the core theme of this competition. I have a few questions, could someone please help me out by explaining in simple terms.</p>\n\n<p>1) Are the tracks (that we try to predict) predefined?\n<a href=\"https://emoyse.web.cern.ch/emoyse/WebEventDisplay/jsdisplay_TrackML.html\">The 3D viewer</a> provided me a better visualisation of the problem, but I don't get a clear picture of what we are trying to predict.</p>\n\n<p>2) What exactly are the 'labels' that we are predicting using the standard DBScan method mentioned in the <a href=\"https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark\">official kernels</a> ?</p>",
      "rawMarkdown": "Hi, \nSorry for a Novice question about the competition. Even after going through the official manual, some good EDA and other kernels I fail to understand the core theme of this competition. I have a few questions, could someone please help me out by explaining in simple terms.\n\n1) Are the tracks (that we try to predict) predefined?\n[The 3D viewer][1] provided me a better visualisation of the problem, but I don't get a clear picture of what we are trying to predict.\n\n2) What exactly are the 'labels' that we are predicting using the standard DBScan method mentioned in the [official kernels][2] ?\n\n\n  [1]: https://emoyse.web.cern.ch/emoyse/WebEventDisplay/jsdisplay_TrackML.html\n  [2]: https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark",
      "votes": null
    },
    {
      "id": "363988",
      "postDate": "07/30/2018 13:27:38",
      "content": "<p>Hi @VikramanK</p>\n\n<p>1) In the training data the track labels for each event are particle_id in the truth.csv files. Multiple hits will have the same particle_id meaning they belong to the same particle or track since if you were to plot the coordinates of each of these hits and draw a line through them it would show the track of the particle. In short, hits are grouped by particle_id to make up a track.</p>\n\n<p>2) There are no truth.csv files for the test data and we do not know how many labels there will be since we don't know how many tracks each event contains. DBSCAN will generate a label (some positive integer) for each input (the hit coordinates being the inputs in this case). e.g. if DBSCAN outputs the number 42 for hits 7,25,49,56,72,89 that means it predicts that these hits form a cluster, group or in this case what we call a track. 42 will be the track_id which is the equivalent of particle_id from the truths.csv file in the training data. </p>",
      "rawMarkdown": "Hi @VikramanK\n\n1) In the training data the track labels for each event are particle_id in the truth.csv files. Multiple hits will have the same particle_id meaning they belong to the same particle or track since if you were to plot the coordinates of each of these hits and draw a line through them it would show the track of the particle. In short, hits are grouped by particle_id to make up a track.\n\n2) There are no truth.csv files for the test data and we do not know how many labels there will be since we don't know how many tracks each event contains. DBSCAN will generate a label (some positive integer) for each input (the hit coordinates being the inputs in this case). e.g. if DBSCAN outputs the number 42 for hits 7,25,49,56,72,89 that means it predicts that these hits form a cluster, group or in this case what we call a track. 42 will be the track_id which is the equivalent of particle_id from the truths.csv file in the training data.",
      "votes": null
    },
    {
      "id": "364312",
      "postDate": "07/31/2018 09:28:11",
      "content": "<p>Hi Jack, \nThank you so much for explaining. After reading your explanation, trackML manual and some EDA, I have a better understanding.  Could you please correct me if the following statements are wrong.\n1) Hits are basically measurements(coordinates) of the particles as detected by the detector. \n2) Assuming we have an instance of a HIT ID from an event, the HIT ID will contain the coordinates of the particle and the truth file will have the corresponding Particle ID. The particle ID is unique for a single particle. A single particle ID can have different HIT IDs corresponding to the different measurements during its trajectory. A single HIT ID can relate to more than one particle ID if two particles' trajectory coincides at the coordinates of this HIT. \n3) We are trying to group is a set of the 3D coordinates, i.e hits that will form a shape of a perturbed helix. </p>\n\n<p>On more clarification, in your example do you mean to say that for hit ID's 7,25,49,56,72,89 the predicted label is 42, meaning that these hit ID's will correspond to a perturbed helix which is equivalent to the path traveled by a particle. What is '42' correspond to ?</p>",
      "rawMarkdown": "Hi Jack, \nThank you so much for explaining. After reading your explanation, trackML manual and some EDA, I have a better understanding.  Could you please correct me if the following statements are wrong.\n1) Hits are basically measurements(coordinates) of the particles as detected by the detector. \n2) Assuming we have an instance of a HIT ID from an event, the HIT ID will contain the coordinates of the particle and the truth file will have the corresponding Particle ID. The particle ID is unique for a single particle. A single particle ID can have different HIT IDs corresponding to the different measurements during its trajectory. A single HIT ID can relate to more than one particle ID if two particles' trajectory coincides at the coordinates of this HIT. \n3) We are trying to group is a set of the 3D coordinates, i.e hits that will form a shape of a perturbed helix. \n\n\nOn more clarification, in your example do you mean to say that for hit ID's 7,25,49,56,72,89 the predicted label is 42, meaning that these hit ID's will correspond to a perturbed helix which is equivalent to the path traveled by a particle. What is '42' correspond to ?",
      "votes": null
    },
    {
      "id": "364454",
      "postDate": "07/31/2018 14:49:03",
      "content": "<p>You're welcome! Everything in 1), 2), 3) looks correct except:</p>\n\n<blockquote>\n  <p>A single HIT ID can relate to more than one particle ID if two\n  particles' trajectory coincides at the coordinates of this HIT</p>\n</blockquote>\n\n<p>To my knowledge a hit can only belong to one particle but hits can be very close together and multiple hits can be detected by one detector cell.</p>\n\n<p>Yes, in my example the coordinates of hits 7,25,49,56,72,89 would correspond to the predicted path of a particle. 42 is the label that DBSCAN generates for the cluster of hits. Basically, you can think of 42 meaning the 42nd cluster found by DBSCAN.</p>",
      "rawMarkdown": "You're welcome! Everything in 1), 2), 3) looks correct except:\n\n&gt; A single HIT ID can relate to more than one particle ID if two\n&gt; particles' trajectory coincides at the coordinates of this HIT\n\nTo my knowledge a hit can only belong to one particle but hits can be very close together and multiple hits can be detected by one detector cell.\n\nYes, in my example the coordinates of hits 7,25,49,56,72,89 would correspond to the predicted path of a particle. 42 is the label that DBSCAN generates for the cluster of hits. Basically, you can think of 42 meaning the 42nd cluster found by DBSCAN.",
      "votes": null
    },
    {
      "id": "364487",
      "postDate": "07/31/2018 16:02:22",
      "content": "<p>Very accurate, @jack vial</p>",
      "rawMarkdown": "Very accurate, @jack vial",
      "votes": null
    },
    {
      "id": "364530",
      "postDate": "07/31/2018 17:30:51",
      "content": "<p>Thanks @David, glad to hear my explanation checks out.</p>",
      "rawMarkdown": "Thanks @David, glad to hear my explanation checks out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 363988,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "07/30/2018 13:27:38",
      "content": "<p>Hi @VikramanK</p>\n\n<p>1) In the training data the track labels for each event are particle_id in the truth.csv files. Multiple hits will have the same particle_id meaning they belong to the same particle or track since if you were to plot the coordinates of each of these hits and draw a line through them it would show the track of the particle. In short, hits are grouped by particle_id to make up a track.</p>\n\n<p>2) There are no truth.csv files for the test data and we do not know how many labels there will be since we don't know how many tracks each event contains. DBSCAN will generate a label (some positive integer) for each input (the hit coordinates being the inputs in this case). e.g. if DBSCAN outputs the number 42 for hits 7,25,49,56,72,89 that means it predicts that these hits form a cluster, group or in this case what we call a track. 42 will be the track_id which is the equivalent of particle_id from the truths.csv file in the training data. </p>",
      "votes": null,
      "replies": [
        {
          "id": 364312,
          "author_name": "vikramank",
          "author_url": "",
          "post_date": "07/31/2018 09:28:11",
          "content": "<p>Hi Jack, \nThank you so much for explaining. After reading your explanation, trackML manual and some EDA, I have a better understanding.  Could you please correct me if the following statements are wrong.\n1) Hits are basically measurements(coordinates) of the particles as detected by the detector. \n2) Assuming we have an instance of a HIT ID from an event, the HIT ID will contain the coordinates of the particle and the truth file will have the corresponding Particle ID. The particle ID is unique for a single particle. A single particle ID can have different HIT IDs corresponding to the different measurements during its trajectory. A single HIT ID can relate to more than one particle ID if two particles' trajectory coincides at the coordinates of this HIT. \n3) We are trying to group is a set of the 3D coordinates, i.e hits that will form a shape of a perturbed helix. </p>\n\n<p>On more clarification, in your example do you mean to say that for hit ID's 7,25,49,56,72,89 the predicted label is 42, meaning that these hit ID's will correspond to a perturbed helix which is equivalent to the path traveled by a particle. What is '42' correspond to ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 364454,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "07/31/2018 14:49:03",
          "content": "<p>You're welcome! Everything in 1), 2), 3) looks correct except:</p>\n\n<blockquote>\n  <p>A single HIT ID can relate to more than one particle ID if two\n  particles' trajectory coincides at the coordinates of this HIT</p>\n</blockquote>\n\n<p>To my knowledge a hit can only belong to one particle but hits can be very close together and multiple hits can be detected by one detector cell.</p>\n\n<p>Yes, in my example the coordinates of hits 7,25,49,56,72,89 would correspond to the predicted path of a particle. 42 is the label that DBSCAN generates for the cluster of hits. Basically, you can think of 42 meaning the 42nd cluster found by DBSCAN.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 364487,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "07/31/2018 16:02:22",
          "content": "<p>Very accurate, @jack vial</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 364530,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "07/31/2018 17:30:51",
          "content": "<p>Thanks @David, glad to hear my explanation checks out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "363951": "Hi, \nSorry for a Novice question about the competition. Even after going through the official manual, some good EDA and other kernels I fail to understand the core theme of this competition. I have a few questions, could someone please help me out by explaining in simple terms.\n\n1) Are the tracks (that we try to predict) predefined?\n[The 3D viewer][1] provided me a better visualisation of the problem, but I don't get a clear picture of what we are trying to predict.\n\n2) What exactly are the 'labels' that we are predicting using the standard DBScan method mentioned in the [official kernels][2] ?\n\n\n  [1]: https://emoyse.web.cern.ch/emoyse/WebEventDisplay/jsdisplay_TrackML.html\n  [2]: https://www.kaggle.com/mikhailhushchyn/dbscan-benchmark",
    "363988": "Hi @VikramanK\n\n1) In the training data the track labels for each event are particle_id in the truth.csv files. Multiple hits will have the same particle_id meaning they belong to the same particle or track since if you were to plot the coordinates of each of these hits and draw a line through them it would show the track of the particle. In short, hits are grouped by particle_id to make up a track.\n\n2) There are no truth.csv files for the test data and we do not know how many labels there will be since we don't know how many tracks each event contains. DBSCAN will generate a label (some positive integer) for each input (the hit coordinates being the inputs in this case). e.g. if DBSCAN outputs the number 42 for hits 7,25,49,56,72,89 that means it predicts that these hits form a cluster, group or in this case what we call a track. 42 will be the track_id which is the equivalent of particle_id from the truths.csv file in the training data.",
    "364312": "Hi Jack, \nThank you so much for explaining. After reading your explanation, trackML manual and some EDA, I have a better understanding.  Could you please correct me if the following statements are wrong.\n1) Hits are basically measurements(coordinates) of the particles as detected by the detector. \n2) Assuming we have an instance of a HIT ID from an event, the HIT ID will contain the coordinates of the particle and the truth file will have the corresponding Particle ID. The particle ID is unique for a single particle. A single particle ID can have different HIT IDs corresponding to the different measurements during its trajectory. A single HIT ID can relate to more than one particle ID if two particles' trajectory coincides at the coordinates of this HIT. \n3) We are trying to group is a set of the 3D coordinates, i.e hits that will form a shape of a perturbed helix. \n\n\nOn more clarification, in your example do you mean to say that for hit ID's 7,25,49,56,72,89 the predicted label is 42, meaning that these hit ID's will correspond to a perturbed helix which is equivalent to the path traveled by a particle. What is '42' correspond to ?",
    "364454": "You're welcome! Everything in 1), 2), 3) looks correct except:\n\n&gt; A single HIT ID can relate to more than one particle ID if two\n&gt; particles' trajectory coincides at the coordinates of this HIT\n\nTo my knowledge a hit can only belong to one particle but hits can be very close together and multiple hits can be detected by one detector cell.\n\nYes, in my example the coordinates of hits 7,25,49,56,72,89 would correspond to the predicted path of a particle. 42 is the label that DBSCAN generates for the cluster of hits. Basically, you can think of 42 meaning the 42nd cluster found by DBSCAN.",
    "364487": "Very accurate, @jack vial",
    "364530": "Thanks @David, glad to hear my explanation checks out."
  },
  "source": "meta"
}