{
  "id": 57020,
  "title": "TrackML erata and other findings",
  "url": "/competitions/trackml-particle-identification/discussion/57020",
  "author_name": "",
  "post_date": "2018-05-18T05:56:29.559411500Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I've started this thread in order to collect errors in the competition <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">document</a> and description. </p>\n\n<p>I've also started this thread in order to collect discovered \"rules\" other than the ones already described in the competition document and description. Let's also say they might event be considered clarifications or findings.</p>\n\n<p>Findings so far:</p>\n\n<p><strong>Errors in the document:</strong></p>\n\n<ul>\n<li><p>Appendix A specifies \"The cells are square pixels of p0 × p1 = 50µm × 50µm.\" - wrong, use \"detectors.csv\" file for correct pitch</p></li>\n<li><p>Appendix A specifies \"For the Long Strip detector, the cell (strip) geometry is also trapezoidal: 120µm × 10.8mm.\" - wrong, use \"detectors.csv\" file for correct pitch</p></li>\n</ul>\n\n<p><strong>Track rules &amp;/ clarifications:</strong></p>\n\n<ul>\n<li><p><em>Data description - Event truth</em> states \"particle_id\" yet it is expected we provide in solution a \"track_id\" - they are the same thing so [\"hit_id\", \"particle_id\"] is the equivalent of saying [\"hit_id\", \"track_id\"]</p></li>\n<li><p><em>Data description - Event truth</em> states \"particle_id [...] A value of 0 means that the hit did not originate from a reconstructible particle, but e.g. from detector noise.\" It might not be clear that a \"weight = 0\" could also be classified as \"noise\" or let's say \"non relevant track\".  So you might want to think about having 2 categories of \"noise\" hits, the rest of hits would have valid tracks.</p></li>\n<li><p>\"track_id\" in the predictions is not based on \"particle_id\" numeric value but will be a number you will pick (generate) to identify the group of hits. “particle_id\" is generated and the numerical value is not the \"truth\", the \"truth\" is the group of hits with the same \"particle_id\". </p></li>\n</ul>\n\n<p><strong>Hypothesis:</strong></p>\n\n<ul>\n<li>The document talks in <em>2.1</em> about an ideal scenario where \"(a) each particle would leave one and only one hit on each layer of the detector\" but it is not the case. I think that with a slight change in the wording we might have something close to that ideal case. That change would be to say \"each particle would leave one and only one hit on each <em>module</em> of the detector\". (demonstration?)</li>\n</ul>\n\n<p>I would do my best to update this based on feedback from others and if I find anything else interesting to add. </p>",
  "messages": [
    {
      "id": "330140",
      "postDate": "05/18/2018 05:56:29",
      "content": "<p>I've started this thread in order to collect errors in the competition <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">document</a> and description. </p>\n\n<p>I've also started this thread in order to collect discovered \"rules\" other than the ones already described in the competition document and description. Let's also say they might event be considered clarifications or findings.</p>\n\n<p>Findings so far:</p>\n\n<p><strong>Errors in the document:</strong></p>\n\n<ul>\n<li><p>Appendix A specifies \"The cells are square pixels of p0 × p1 = 50µm × 50µm.\" - wrong, use \"detectors.csv\" file for correct pitch</p></li>\n<li><p>Appendix A specifies \"For the Long Strip detector, the cell (strip) geometry is also trapezoidal: 120µm × 10.8mm.\" - wrong, use \"detectors.csv\" file for correct pitch</p></li>\n</ul>\n\n<p><strong>Track rules &amp;/ clarifications:</strong></p>\n\n<ul>\n<li><p><em>Data description - Event truth</em> states \"particle_id\" yet it is expected we provide in solution a \"track_id\" - they are the same thing so [\"hit_id\", \"particle_id\"] is the equivalent of saying [\"hit_id\", \"track_id\"]</p></li>\n<li><p><em>Data description - Event truth</em> states \"particle_id [...] A value of 0 means that the hit did not originate from a reconstructible particle, but e.g. from detector noise.\" It might not be clear that a \"weight = 0\" could also be classified as \"noise\" or let's say \"non relevant track\".  So you might want to think about having 2 categories of \"noise\" hits, the rest of hits would have valid tracks.</p></li>\n<li><p>\"track_id\" in the predictions is not based on \"particle_id\" numeric value but will be a number you will pick (generate) to identify the group of hits. “particle_id\" is generated and the numerical value is not the \"truth\", the \"truth\" is the group of hits with the same \"particle_id\". </p></li>\n</ul>\n\n<p><strong>Hypothesis:</strong></p>\n\n<ul>\n<li>The document talks in <em>2.1</em> about an ideal scenario where \"(a) each particle would leave one and only one hit on each layer of the detector\" but it is not the case. I think that with a slight change in the wording we might have something close to that ideal case. That change would be to say \"each particle would leave one and only one hit on each <em>module</em> of the detector\". (demonstration?)</li>\n</ul>\n\n<p>I would do my best to update this based on feedback from others and if I find anything else interesting to add. </p>",
      "rawMarkdown": "I've started this thread in order to collect errors in the competition [document][1] and description. \n\nI've also started this thread in order to collect discovered \"rules\" other than the ones already described in the competition document and description. Let's also say they might event be considered clarifications or findings.\n\nFindings so far:\n\n**Errors in the document:**\n\n- Appendix A specifies \"The cells are square pixels of p0 × p1 = 50µm × 50µm.\" - wrong, use \"detectors.csv\" file for correct pitch\n\n- Appendix A specifies \"For the Long Strip detector, the cell (strip) geometry is also trapezoidal: 120µm × 10.8mm.\" - wrong, use \"detectors.csv\" file for correct pitch\n\n\n**Track rules &amp;/ clarifications:**\n\n- _Data description - Event truth_ states \"particle_id\" yet it is expected we provide in solution a \"track_id\" - they are the same thing so [\"hit_id\", \"particle_id\"] is the equivalent of saying [\"hit_id\", \"track_id\"]\n\n- _Data description - Event truth_ states \"particle_id [...] A value of 0 means that the hit did not originate from a reconstructible particle, but e.g. from detector noise.\" It might not be clear that a \"weight = 0\" could also be classified as \"noise\" or let's say \"non relevant track\".  So you might want to think about having 2 categories of \"noise\" hits, the rest of hits would have valid tracks.\n\n- \"track_id\" in the predictions is not based on \"particle_id\" numeric value but will be a number you will pick (generate) to identify the group of hits. “particle_id\" is generated and the numerical value is not the \"truth\", the \"truth\" is the group of hits with the same \"particle_id\". \n\n**Hypothesis:**\n\n- The document talks in _2.1_ about an ideal scenario where \"(a) each particle would leave one and only one hit on each layer of the detector\" but it is not the case. I think that with a slight change in the wording we might have something close to that ideal case. That change would be to say \"each particle would leave one and only one hit on each _module_ of the detector\". (demonstration?)\n\n\n\nI would do my best to update this based on feedback from others and if I find anything else interesting to add. \n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf",
      "votes": null
    },
    {
      "id": "330274",
      "postDate": "05/18/2018 13:04:25",
      "content": "<p>Thanks for this. We are actually updating the documents.  One important point: we have very déliberately decided to use «track_id&nbsp;» and not «&nbsp;particle_id&nbsp;» to make it as clear as possible that we are not interested by the value of track_id</p>",
      "rawMarkdown": "Thanks for this. We are actually updating the documents.  One important point: we have very déliberately decided to use «track_id&nbsp;» and not «&nbsp;particle_id&nbsp;» to make it as clear as possible that we are not interested by the value of track_id",
      "votes": null
    },
    {
      "id": "330279",
      "postDate": "05/18/2018 13:17:08",
      "content": "<p>@David I think you mean you are interested in the value of \"track_id\", not from the point of view of \"particle\" but from the point of view of grouping at least 50% of the hits from a track generated by a particle.</p>\n\n<p>My post tries to point out that \"particle_id\" is the same as saying \"track_id\". It's just that at train time there's way more info available from a \"particle\" perspective. </p>",
      "rawMarkdown": "David I think you mean you are interested in the value of \"track_id\", not from the point of view of \"particle\" but from the point of view of grouping at least 50% of the hits from a track generated by a particle.\n\nMy post tries to point out that \"particle_id\" is the same as saying \"track_id\". It's just that at train time there's way more info available from a \"particle\" perspective.",
      "votes": null
    },
    {
      "id": "330281",
      "postDate": "05/18/2018 13:25:10",
      "content": "<p>That s correct, but it would have been confusing to use the same name.  </p>",
      "rawMarkdown": "That s correct, but it would have been confusing to use the same name.",
      "votes": null
    },
    {
      "id": "330293",
      "postDate": "05/18/2018 13:55:21",
      "content": "<p>@David I've updated the main post in an attempt to clarify to others what you mean by \"not interested in track_id value\". </p>",
      "rawMarkdown": "David I've updated the main post in an attempt to clarify to others what you mean by \"not interested in track_id value\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 330274,
      "author_name": "droussea",
      "author_url": "",
      "post_date": "05/18/2018 13:04:25",
      "content": "<p>Thanks for this. We are actually updating the documents.  One important point: we have very déliberately decided to use «track_id&nbsp;» and not «&nbsp;particle_id&nbsp;» to make it as clear as possible that we are not interested by the value of track_id</p>",
      "votes": null,
      "replies": [
        {
          "id": 330279,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/18/2018 13:17:08",
          "content": "<p>@David I think you mean you are interested in the value of \"track_id\", not from the point of view of \"particle\" but from the point of view of grouping at least 50% of the hits from a track generated by a particle.</p>\n\n<p>My post tries to point out that \"particle_id\" is the same as saying \"track_id\". It's just that at train time there's way more info available from a \"particle\" perspective. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 330281,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "05/18/2018 13:25:10",
          "content": "<p>That s correct, but it would have been confusing to use the same name.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 330293,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/18/2018 13:55:21",
          "content": "<p>@David I've updated the main post in an attempt to clarify to others what you mean by \"not interested in track_id value\". </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "330140": "I've started this thread in order to collect errors in the competition [document][1] and description. \n\nI've also started this thread in order to collect discovered \"rules\" other than the ones already described in the competition document and description. Let's also say they might event be considered clarifications or findings.\n\nFindings so far:\n\n**Errors in the document:**\n\n- Appendix A specifies \"The cells are square pixels of p0 × p1 = 50µm × 50µm.\" - wrong, use \"detectors.csv\" file for correct pitch\n\n- Appendix A specifies \"For the Long Strip detector, the cell (strip) geometry is also trapezoidal: 120µm × 10.8mm.\" - wrong, use \"detectors.csv\" file for correct pitch\n\n\n**Track rules &amp;/ clarifications:**\n\n- _Data description - Event truth_ states \"particle_id\" yet it is expected we provide in solution a \"track_id\" - they are the same thing so [\"hit_id\", \"particle_id\"] is the equivalent of saying [\"hit_id\", \"track_id\"]\n\n- _Data description - Event truth_ states \"particle_id [...] A value of 0 means that the hit did not originate from a reconstructible particle, but e.g. from detector noise.\" It might not be clear that a \"weight = 0\" could also be classified as \"noise\" or let's say \"non relevant track\".  So you might want to think about having 2 categories of \"noise\" hits, the rest of hits would have valid tracks.\n\n- \"track_id\" in the predictions is not based on \"particle_id\" numeric value but will be a number you will pick (generate) to identify the group of hits. “particle_id\" is generated and the numerical value is not the \"truth\", the \"truth\" is the group of hits with the same \"particle_id\". \n\n**Hypothesis:**\n\n- The document talks in _2.1_ about an ideal scenario where \"(a) each particle would leave one and only one hit on each layer of the detector\" but it is not the case. I think that with a slight change in the wording we might have something close to that ideal case. That change would be to say \"each particle would leave one and only one hit on each _module_ of the detector\". (demonstration?)\n\n\n\nI would do my best to update this based on feedback from others and if I find anything else interesting to add. \n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf",
    "330274": "Thanks for this. We are actually updating the documents.  One important point: we have very déliberately decided to use «track_id&nbsp;» and not «&nbsp;particle_id&nbsp;» to make it as clear as possible that we are not interested by the value of track_id",
    "330279": "David I think you mean you are interested in the value of \"track_id\", not from the point of view of \"particle\" but from the point of view of grouping at least 50% of the hits from a track generated by a particle.\n\nMy post tries to point out that \"particle_id\" is the same as saying \"track_id\". It's just that at train time there's way more info available from a \"particle\" perspective.",
    "330281": "That s correct, but it would have been confusing to use the same name.",
    "330293": "David I've updated the main post in an attempt to clarify to others what you mean by \"not interested in track_id value\"."
  },
  "source": "meta"
}