{
  "id": 55796,
  "title": "`track_id` to `hit_id` Mapping Assumptions",
  "url": "/competitions/trackml-particle-identification/discussion/55796",
  "author_name": "",
  "post_date": "2018-05-01T20:42:48.980353Z",
  "votes": null,
  "comment_count": 12,
  "views": 0,
  "content": "<p>In this competition we are mapping every <code>track_id</code> to one <code>hit_id</code>.</p>\n\n<p>To do this, we need to infer the particle's trajectory and generate our own <code>track_id</code> for every track we reconstruct.  Then we map every <code>track_id</code> to one <code>hit_id</code>.  Since every <code>hit_id</code> can be mapped to one particle in the given truth data (or validation data), we are basically trying to identify the number of hits for each particle.  This seems to be the only information you can get from our submissions.</p>\n\n<p>To competition organizers: Is my reasoning valid?</p>",
  "messages": [
    {
      "id": "321744",
      "postDate": "05/01/2018 20:42:48",
      "content": "<p>In this competition we are mapping every <code>track_id</code> to one <code>hit_id</code>.</p>\n\n<p>To do this, we need to infer the particle's trajectory and generate our own <code>track_id</code> for every track we reconstruct.  Then we map every <code>track_id</code> to one <code>hit_id</code>.  Since every <code>hit_id</code> can be mapped to one particle in the given truth data (or validation data), we are basically trying to identify the number of hits for each particle.  This seems to be the only information you can get from our submissions.</p>\n\n<p>To competition organizers: Is my reasoning valid?</p>",
      "rawMarkdown": "In this competition we are mapping every `track_id` to one `hit_id`.\n\nTo do this, we need to infer the particle's trajectory and generate our own `track_id` for every track we reconstruct.  Then we map every `track_id` to one `hit_id`.  Since every `hit_id` can be mapped to one particle in the given truth data (or validation data), we are basically trying to identify the number of hits for each particle.  This seems to be the only information you can get from our submissions.\n\nTo competition organizers: Is my reasoning valid?",
      "votes": null
    },
    {
      "id": "321750",
      "postDate": "05/01/2018 20:52:29",
      "content": "<p>Almost, we’re in fact interested “which” hits stem from each particle, not only “how many”. This is why we ask for ‘hit_id’ s associated to ‘track_id’ s - where for latter you, of course, can take your own numbering. </p>",
      "rawMarkdown": "Almost, we’re in fact interested “which” hits stem from each particle, not only “how many”. This is why we ask for ‘hit_id’ s associated to ‘track_id’ s - where for latter you, of course, can take your own numbering.",
      "votes": null
    },
    {
      "id": "321760",
      "postDate": "05/01/2018 21:09:16",
      "content": "<p>Thanks for the quick answer.</p>\n\n<p>How can you link a particle to my generated <code>track_id</code> value? so you can later map it a <code>hit_id</code> with my submission.</p>",
      "rawMarkdown": "Thanks for the quick answer.\n\nHow can you link a particle to my generated `track_id` value? so you can later map it a `hit_id` with my submission.",
      "votes": null
    },
    {
      "id": "321761",
      "postDate": "05/01/2018 21:10:34",
      "content": "<p>To be completely clear, track_id is an arbitrary non negative integer. All what we are interested in is that all hit_id associated to the same track_id are together. It is like sorting beans in piles, and then labelling the piles 1 2 3 4 or 123 456 234 20, does not make any difference provided the labels are distinct integer.</p>",
      "rawMarkdown": "To be completely clear, track_id is an arbitrary non negative integer. All what we are interested in is that all hit_id associated to the same track_id are together. It is like sorting beans in piles, and then labelling the piles 1 2 3 4 or 123 456 234 20, does not make any difference provided the labels are distinct integer.",
      "votes": null
    },
    {
      "id": "321764",
      "postDate": "05/01/2018 21:13:14",
      "content": "<p>By the way : we ask every hit_id to be associated to a track_id, but it is completely fine to define a garbage track (with track_id 0 maybe) will all the hits your algorithm could not assigned. This garbage track will contribute zero to the score of course, but this will make a valid contribution.</p>",
      "rawMarkdown": "By the way : we ask every hit_id to be associated to a track_id, but it is completely fine to define a garbage track (with track_id 0 maybe) will all the hits your algorithm could not assigned. This garbage track will contribute zero to the score of course, but this will make a valid contribution.",
      "votes": null
    },
    {
      "id": "321765",
      "postDate": "05/01/2018 21:14:21",
      "content": "<p>Good analogy, David.  That also means the only meaningful information we get is how big each pile is.  Every bean (<code>track_id</code>) in non-distinguishable from the rest.</p>",
      "rawMarkdown": "Good analogy, David.  That also means the only meaningful information we get is how big each pile is.  Every bean (`track_id`) in non-distinguishable from the rest.",
      "votes": null
    },
    {
      "id": "321797",
      "postDate": "05/01/2018 23:03:28",
      "content": "<p>As far as I read you need to submit a csv submission file with the following information:</p>\n\n<pre><code>event_id, hit_id, track_id\n</code></pre>\n\n<p>Where each hit <strong>of an event</strong> must belong to exactly one track_id (track_id is unique for hits of an event). And the track_id is simply some arbitrary id used for the purpose of grouping your hits. So one track_id can be associated with multiple hit_ids (if the track contains more than one data point), but not to more than one hit within the same event.</p>\n\n<p>See <a href=\"https://www.kaggle.com/c/trackml-particle-identification/data\">Dataset submission information</a></p>\n\n<blockquote>\n  <p>The submission file must associate each hit in each event to one and\n  only one reconstructed particle track. The reconstructed tracks must\n  be uniquely identified only within each event. Participants are\n  advised to compress the submission file (with zip, bzip2, gzip) before\n  submission to the Kaggle site.</p>\n</blockquote>",
      "rawMarkdown": "As far as I read you need to submit a csv submission file with the following information:\n\n    event_id, hit_id, track_id\n\nWhere each hit **of an event** must belong to exactly one track_id (track_id is unique for hits of an event). And the track_id is simply some arbitrary id used for the purpose of grouping your hits. So one track_id can be associated with multiple hit_ids (if the track contains more than one data point), but not to more than one hit within the same event.\n\nSee [Dataset submission information][1]\n\n&gt; The submission file must associate each hit in each event to one and\n&gt; only one reconstructed particle track. The reconstructed tracks must\n&gt; be uniquely identified only within each event. Participants are\n&gt; advised to compress the submission file (with zip, bzip2, gzip) before\n&gt; submission to the Kaggle site.\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/data",
      "votes": null
    },
    {
      "id": "321984",
      "postDate": "05/02/2018 08:58:15",
      "content": "<blockquote>\n  <p>How can you link a particle to my generated track_id value? so you can\n  later map it a hit_id with my submission.</p>\n</blockquote>\n\n<p>if I understand correctly, you ask how we map the reconstructed tracks to the ground truth particles. This is a very good question. In the evaluation page: </p>\n\n<blockquote>\n  <p>for a given track, the matching particle is the one to which the\n  absolute majority (strictly more that 50%) of the track points belong.</p>\n</blockquote>\n\n<p>See the evaluation page, or section \"The score\" in <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf</a> for more and important details.</p>",
      "rawMarkdown": "&gt; How can you link a particle to my generated track_id value? so you can\n&gt; later map it a hit_id with my submission.\n\nif I understand correctly, you ask how we map the reconstructed tracks to the ground truth particles. This is a very good question. In the evaluation page: \n\n&gt; for a given track, the matching particle is the one to which the\n&gt; absolute majority (strictly more that 50%) of the track points belong.\n\n\nSee the evaluation page, or section \"The score\" in https://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf for more and important details.",
      "votes": null
    },
    {
      "id": "325851",
      "postDate": "05/09/2018 00:49:36",
      "content": "<p>I guess in the analogy every bean means - hit_id</p>",
      "rawMarkdown": "I guess in the analogy every bean means - hit_id",
      "votes": null
    },
    {
      "id": "326103",
      "postDate": "05/09/2018 09:32:52",
      "content": "<p>I think you are wrong, if we are going to map single hit_id to single track_id then there is no competition, because we can just define track_id equal to hit_id and score 100%.\nI understand that we need to group multiple hits into the tracks. So multiple hit per one track.</p>",
      "rawMarkdown": "I think you are wrong, if we are going to map single hit_id to single track_id then there is no competition, because we can just define track_id equal to hit_id and score 100%.\nI understand that we need to group multiple hits into the tracks. So multiple hit per one track.",
      "votes": null
    },
    {
      "id": "326149",
      "postDate": "05/09/2018 11:04:42",
      "content": "<p>Yes, so you make piles of hit_id, and then give an arbitrary  number (track_id) to each pile.</p>",
      "rawMarkdown": "Yes, so you make piles of hit_id, and then give an arbitrary  number (track_id) to each pile.",
      "votes": null
    },
    {
      "id": "326151",
      "postDate": "05/09/2018 11:06:08",
      "content": "<p>Right (But Wesam as realized this long ago)</p>",
      "rawMarkdown": "Right (But Wesam as realized this long ago)",
      "votes": null
    },
    {
      "id": "326154",
      "postDate": "05/09/2018 11:10:25",
      "content": "<p>just started with this competition so I read all discussions :)</p>",
      "rawMarkdown": "just started with this competition so I read all discussions :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 321750,
      "author_name": "asalzburger",
      "author_url": "",
      "post_date": "05/01/2018 20:52:29",
      "content": "<p>Almost, we’re in fact interested “which” hits stem from each particle, not only “how many”. This is why we ask for ‘hit_id’ s associated to ‘track_id’ s - where for latter you, of course, can take your own numbering. </p>",
      "votes": null,
      "replies": [
        {
          "id": 321760,
          "author_name": "wesamelshamy",
          "author_url": "",
          "post_date": "05/01/2018 21:09:16",
          "content": "<p>Thanks for the quick answer.</p>\n\n<p>How can you link a particle to my generated <code>track_id</code> value? so you can later map it a <code>hit_id</code> with my submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321797,
          "author_name": "alexec",
          "author_url": "",
          "post_date": "05/01/2018 23:03:28",
          "content": "<p>As far as I read you need to submit a csv submission file with the following information:</p>\n\n<pre><code>event_id, hit_id, track_id\n</code></pre>\n\n<p>Where each hit <strong>of an event</strong> must belong to exactly one track_id (track_id is unique for hits of an event). And the track_id is simply some arbitrary id used for the purpose of grouping your hits. So one track_id can be associated with multiple hit_ids (if the track contains more than one data point), but not to more than one hit within the same event.</p>\n\n<p>See <a href=\"https://www.kaggle.com/c/trackml-particle-identification/data\">Dataset submission information</a></p>\n\n<blockquote>\n  <p>The submission file must associate each hit in each event to one and\n  only one reconstructed particle track. The reconstructed tracks must\n  be uniquely identified only within each event. Participants are\n  advised to compress the submission file (with zip, bzip2, gzip) before\n  submission to the Kaggle site.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321984,
          "author_name": "cecilegermain",
          "author_url": "",
          "post_date": "05/02/2018 08:58:15",
          "content": "<blockquote>\n  <p>How can you link a particle to my generated track_id value? so you can\n  later map it a hit_id with my submission.</p>\n</blockquote>\n\n<p>if I understand correctly, you ask how we map the reconstructed tracks to the ground truth particles. This is a very good question. In the evaluation page: </p>\n\n<blockquote>\n  <p>for a given track, the matching particle is the one to which the\n  absolute majority (strictly more that 50%) of the track points belong.</p>\n</blockquote>\n\n<p>See the evaluation page, or section \"The score\" in <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf</a> for more and important details.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321761,
      "author_name": "droussea",
      "author_url": "",
      "post_date": "05/01/2018 21:10:34",
      "content": "<p>To be completely clear, track_id is an arbitrary non negative integer. All what we are interested in is that all hit_id associated to the same track_id are together. It is like sorting beans in piles, and then labelling the piles 1 2 3 4 or 123 456 234 20, does not make any difference provided the labels are distinct integer.</p>",
      "votes": null,
      "replies": [
        {
          "id": 321765,
          "author_name": "wesamelshamy",
          "author_url": "",
          "post_date": "05/01/2018 21:14:21",
          "content": "<p>Good analogy, David.  That also means the only meaningful information we get is how big each pile is.  Every bean (<code>track_id</code>) in non-distinguishable from the rest.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325851,
          "author_name": "superq8",
          "author_url": "",
          "post_date": "05/09/2018 00:49:36",
          "content": "<p>I guess in the analogy every bean means - hit_id</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326149,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "05/09/2018 11:04:42",
          "content": "<p>Yes, so you make piles of hit_id, and then give an arbitrary  number (track_id) to each pile.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321764,
      "author_name": "droussea",
      "author_url": "",
      "post_date": "05/01/2018 21:13:14",
      "content": "<p>By the way : we ask every hit_id to be associated to a track_id, but it is completely fine to define a garbage track (with track_id 0 maybe) will all the hits your algorithm could not assigned. This garbage track will contribute zero to the score of course, but this will make a valid contribution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326103,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "05/09/2018 09:32:52",
      "content": "<p>I think you are wrong, if we are going to map single hit_id to single track_id then there is no competition, because we can just define track_id equal to hit_id and score 100%.\nI understand that we need to group multiple hits into the tracks. So multiple hit per one track.</p>",
      "votes": null,
      "replies": [
        {
          "id": 326151,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "05/09/2018 11:06:08",
          "content": "<p>Right (But Wesam as realized this long ago)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326154,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "05/09/2018 11:10:25",
          "content": "<p>just started with this competition so I read all discussions :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "321744": "In this competition we are mapping every `track_id` to one `hit_id`.\n\nTo do this, we need to infer the particle's trajectory and generate our own `track_id` for every track we reconstruct.  Then we map every `track_id` to one `hit_id`.  Since every `hit_id` can be mapped to one particle in the given truth data (or validation data), we are basically trying to identify the number of hits for each particle.  This seems to be the only information you can get from our submissions.\n\nTo competition organizers: Is my reasoning valid?",
    "321750": "Almost, we’re in fact interested “which” hits stem from each particle, not only “how many”. This is why we ask for ‘hit_id’ s associated to ‘track_id’ s - where for latter you, of course, can take your own numbering.",
    "321760": "Thanks for the quick answer.\n\nHow can you link a particle to my generated `track_id` value? so you can later map it a `hit_id` with my submission.",
    "321761": "To be completely clear, track_id is an arbitrary non negative integer. All what we are interested in is that all hit_id associated to the same track_id are together. It is like sorting beans in piles, and then labelling the piles 1 2 3 4 or 123 456 234 20, does not make any difference provided the labels are distinct integer.",
    "321764": "By the way : we ask every hit_id to be associated to a track_id, but it is completely fine to define a garbage track (with track_id 0 maybe) will all the hits your algorithm could not assigned. This garbage track will contribute zero to the score of course, but this will make a valid contribution.",
    "321765": "Good analogy, David.  That also means the only meaningful information we get is how big each pile is.  Every bean (`track_id`) in non-distinguishable from the rest.",
    "321797": "As far as I read you need to submit a csv submission file with the following information:\n\n    event_id, hit_id, track_id\n\nWhere each hit **of an event** must belong to exactly one track_id (track_id is unique for hits of an event). And the track_id is simply some arbitrary id used for the purpose of grouping your hits. So one track_id can be associated with multiple hit_ids (if the track contains more than one data point), but not to more than one hit within the same event.\n\nSee [Dataset submission information][1]\n\n&gt; The submission file must associate each hit in each event to one and\n&gt; only one reconstructed particle track. The reconstructed tracks must\n&gt; be uniquely identified only within each event. Participants are\n&gt; advised to compress the submission file (with zip, bzip2, gzip) before\n&gt; submission to the Kaggle site.\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/data",
    "321984": "&gt; How can you link a particle to my generated track_id value? so you can\n&gt; later map it a hit_id with my submission.\n\nif I understand correctly, you ask how we map the reconstructed tracks to the ground truth particles. This is a very good question. In the evaluation page: \n\n&gt; for a given track, the matching particle is the one to which the\n&gt; absolute majority (strictly more that 50%) of the track points belong.\n\n\nSee the evaluation page, or section \"The score\" in https://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf for more and important details.",
    "325851": "I guess in the analogy every bean means - hit_id",
    "326103": "I think you are wrong, if we are going to map single hit_id to single track_id then there is no competition, because we can just define track_id equal to hit_id and score 100%.\nI understand that we need to group multiple hits into the tracks. So multiple hit per one track.",
    "326149": "Yes, so you make piles of hit_id, and then give an arbitrary  number (track_id) to each pile.",
    "326151": "Right (But Wesam as realized this long ago)",
    "326154": "just started with this competition so I read all discussions :)"
  },
  "source": "meta"
}