{
  "id": 56684,
  "title": "Calculating the cell coordinates",
  "url": "/competitions/trackml-particle-identification/discussion/56684",
  "author_name": "",
  "post_date": "2018-05-13T12:36:22.221799400Z",
  "votes": null,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Dear organizers,</p>\n\n<p>Firstly, thanks a lot for making such a competition possible!</p>\n\n<p>Now, I'm trying to get the xyz coordinates of the cells using ch0 and ch1 parameters.\n<a href=\"https://www.kaggle.com/silikhon/cellposition\">Here's the code I wrote for that</a>. What bugs me is that the mean difference between the provided hit xyz coordinates and the calculated coords of the cells involved (weighted by the cell value, assuming it's proportional to the energy deposition within that cell) has some weird distribution:\n<img src=\"https://cernbox.cern.ch/index.php/s/A6cDax9t1nMOyyU/download\" alt=\"Weighted mean distance between the calculated cell center and hit position\">\n(This is a single-event distribution, evt # 000001100 in particular - but I've tried various events and the distribution looks similarly)</p>\n\n<p>Does this simply mean that hit position is defined in a more complicated way than just cell position mean, or is there a problem in my cell position calculation?</p>\n\n<p>Would appreciate any help. Thanks!</p>",
  "messages": [
    {
      "id": "328129",
      "postDate": "05/13/2018 12:36:22",
      "content": "<p>Dear organizers,</p>\n\n<p>Firstly, thanks a lot for making such a competition possible!</p>\n\n<p>Now, I'm trying to get the xyz coordinates of the cells using ch0 and ch1 parameters.\n<a href=\"https://www.kaggle.com/silikhon/cellposition\">Here's the code I wrote for that</a>. What bugs me is that the mean difference between the provided hit xyz coordinates and the calculated coords of the cells involved (weighted by the cell value, assuming it's proportional to the energy deposition within that cell) has some weird distribution:\n<img src=\"https://cernbox.cern.ch/index.php/s/A6cDax9t1nMOyyU/download\" alt=\"Weighted mean distance between the calculated cell center and hit position\">\n(This is a single-event distribution, evt # 000001100 in particular - but I've tried various events and the distribution looks similarly)</p>\n\n<p>Does this simply mean that hit position is defined in a more complicated way than just cell position mean, or is there a problem in my cell position calculation?</p>\n\n<p>Would appreciate any help. Thanks!</p>",
      "rawMarkdown": "Dear organizers,\n\nFirstly, thanks a lot for making such a competition possible!\n\nNow, I'm trying to get the xyz coordinates of the cells using ch0 and ch1 parameters.\n[Here's the code I wrote for that][1]. What bugs me is that the mean difference between the provided hit xyz coordinates and the calculated coords of the cells involved (weighted by the cell value, assuming it's proportional to the energy deposition within that cell) has some weird distribution:\n![Weighted mean distance between the calculated cell center and hit position][2]\n(This is a single-event distribution, evt # 000001100 in particular - but I've tried various events and the distribution looks similarly)\n\nDoes this simply mean that hit position is defined in a more complicated way than just cell position mean, or is there a problem in my cell position calculation?\n\nWould appreciate any help. Thanks!\n\n  [1]: https://www.kaggle.com/silikhon/cellposition\n  [2]: https://cernbox.cern.ch/index.php/s/A6cDax9t1nMOyyU/download",
      "votes": null
    },
    {
      "id": "328677",
      "postDate": "05/14/2018 21:58:50",
      "content": "<p><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf</a>.  See Appendix A. I would suggest trying a code where you calculate the estimated (X, Y) by doing first (m0,m1) and apply rot after that - not before.  </p>\n\n<p>I've also noticed you're ignoring rot_xw, rot_yw and rot_zw - the cell doesn't have a thickness of 0! It's true that the \"ch3\" is not provide, but maybe the W position is -module_T? So W can alter the value of X, Y, Z. </p>\n\n<p>In theory the distance should be then 0.</p>",
      "rawMarkdown": "https://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf.  See Appendix A. I would suggest trying a code where you calculate the estimated (X, Y) by doing first (m0,m1) and apply rot after that - not before.  \n\nI've also noticed you're ignoring rot_xw, rot_yw and rot_zw - the cell doesn't have a thickness of 0! It's true that the \"ch3\" is not provide, but maybe the W position is -module_T? So W can alter the value of X, Y, Z. \n\nIn theory the distance should be then 0.",
      "votes": null
    },
    {
      "id": "328682",
      "postDate": "05/14/2018 22:36:31",
      "content": "<p>Thanks for the suggestion.</p>\n\n<p>I doubt that order of rotation and averaging matters here: averaging is a linear operation and it commutes with rotation - if only it has some effects of finite numeric precision, I'll try and check that later.\nWhat does matter though is that I average the <em>distance</em> between cell centers and provided hit positions, rather than make an average of dx, dy and dz separately. I've tried this out and got the following dist:\n<img src=\"https://cernbox.cern.ch/index.php/s/MOl03nMQ8fr1XO4/download\" alt=\"mean dist\"></p>\n\n<p>I'll try swapping the rotation and averaging to see if it goes to 0 exactly.\nNote that smallest module width is 0.150 which is much bigger then the discrepancy I'm having here, so it can't be it.</p>\n\n<p>Thanks for pointing me to the right spot!</p>",
      "rawMarkdown": "Thanks for the suggestion.\n\nI doubt that order of rotation and averaging matters here: averaging is a linear operation and it commutes with rotation - if only it has some effects of finite numeric precision, I'll try and check that later.\nWhat does matter though is that I average the *distance* between cell centers and provided hit positions, rather than make an average of dx, dy and dz separately. I've tried this out and got the following dist:\n![mean dist][1]\n\nI'll try swapping the rotation and averaging to see if it goes to 0 exactly.\nNote that smallest module width is 0.150 which is much bigger then the discrepancy I'm having here, so it can't be it.\n\nThanks for pointing me to the right spot!\n\n  [1]: https://cernbox.cern.ch/index.php/s/MOl03nMQ8fr1XO4/download",
      "votes": null
    },
    {
      "id": "328789",
      "postDate": "05/15/2018 04:49:14",
      "content": "<p>I think you're wrong about distances. Something else I've noticed in your code is that you don't take into account the fact that some modules are trapezoid and assume the number of cells on the smaller base to be the same as on the bigger base and it might not be the case.</p>",
      "rawMarkdown": "I think you're wrong about distances. Something else I've noticed in your code is that you don't take into account the fact that some modules are trapezoid and assume the number of cells on the smaller base to be the same as on the bigger base and it might not be the case.",
      "votes": null
    },
    {
      "id": "328799",
      "postDate": "05/15/2018 05:09:15",
      "content": "<p>Actually I still can't get my head around distances. You are saying that the sum of distance between a point and other points is the same as distance of that point to the point resulted from the sum of points. I don't think that's true. Let's take 0 as our point and -1 and 1 the other points. Distance from 0 to -1 is 1, to 1 is 1, sum is 2. Distance to sum of points (-1 + 1) is 0. </p>",
      "rawMarkdown": "Actually I still can't get my head around distances. You are saying that the sum of distance between a point and other points is the same as distance of that point to the point resulted from the sum of points. I don't think that's true. Let's take 0 as our point and -1 and 1 the other points. Distance from 0 to -1 is 1, to 1 is 1, sum is 2. Distance to sum of points (-1 + 1) is 0.",
      "votes": null
    },
    {
      "id": "328816",
      "postDate": "05/15/2018 06:11:14",
      "content": "<p>Regarding the trapezoid shape, I'm following this comment <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/56266#327664\">https://www.kaggle.com/c/trackml-particle-identification/discussion/56266#327664</a> by Andreas, stating that grid is still rectangular for simplicity.\nI'm assuming this means the cells are labeled as if they were in a rectangular module of shape (2 * module_hv,  2 * module_maxhu), and just indexes of cells cut off by the trapezoidal corners never occur. That actually seems to be true. Here's the same plot as before just for the trapezoidal modules:\n<img src=\"https://cernbox.cern.ch/index.php/s/SweosZKPJozLkuQ/download\" alt=\"mean dist for trapezoid\">\nNote that smallest possible step between cells in trapezoid modules is 0.08mm, which is much bigger than the diff I get on this hist. If I missed cell position by one or more steps, you'd see it in this figure as an entry of 0.08 or greater.</p>\n\n<p>Now on your second comment:</p>\n\n<pre><code>Let's take 0 as our point and -1 and 1 the other points. Distance from 0 to -1 is 1, to 1 is 1, sum is 2. Distance to sum of points (-1 + 1) is 0.\n</code></pre>\n\n<p>this was exactly the problem with the code in the kernel. As I said before, once you change that to average of dx, dy and dz separately, you get the distribution as in my last comment. Applying that to your example: dx between -1 and 0 is -1, between 1 and 0 is 1 which makes sum(dx) = -1 + 1 = 0, which is same as dx between (-1 + 1) and 0.</p>\n\n<p>P.S.\nI haven't updated the kernel yet.</p>",
      "rawMarkdown": "Regarding the trapezoid shape, I'm following this comment https://www.kaggle.com/c/trackml-particle-identification/discussion/56266#327664 by Andreas, stating that grid is still rectangular for simplicity.\nI'm assuming this means the cells are labeled as if they were in a rectangular module of shape (2 * module_hv,  2 * module_maxhu), and just indexes of cells cut off by the trapezoidal corners never occur. That actually seems to be true. Here's the same plot as before just for the trapezoidal modules:\n![mean dist for trapezoid][1]\nNote that smallest possible step between cells in trapezoid modules is 0.08mm, which is much bigger than the diff I get on this hist. If I missed cell position by one or more steps, you'd see it in this figure as an entry of 0.08 or greater.\n\nNow on your second comment:\n\n    Let's take 0 as our point and -1 and 1 the other points. Distance from 0 to -1 is 1, to 1 is 1, sum is 2. Distance to sum of points (-1 + 1) is 0.\n\nthis was exactly the problem with the code in the kernel. As I said before, once you change that to average of dx, dy and dz separately, you get the distribution as in my last comment. Applying that to your example: dx between -1 and 0 is -1, between 1 and 0 is 1 which makes sum(dx) = -1 + 1 = 0, which is same as dx between (-1 + 1) and 0.\n\nP.S.\nI haven't updated the kernel yet.\n\n  [1]: https://cernbox.cern.ch/index.php/s/SweosZKPJozLkuQ/download",
      "votes": null
    },
    {
      "id": "328824",
      "postDate": "05/15/2018 06:31:59",
      "content": "<p>Here is how to properly calculate it:</p>\n\n<blockquote>\n  <p>hits, cells, particles, truth = load_event('train_100_events/event000001000')</p>\n  \n  <p>detectors = pd.read_csv(\"detectors.csv\")</p>\n  \n  <p>hits_aug = hits.merge(detectors, on=[\"volume_id\", \"layer_id\", \"module_id\"]).merge(cells, on=\"hit_id\")</p>\n  \n  <p>hits_aug[\"u\"] = ((hits_aug.ch0 + 0.5) * hits_aug.pitch_u - (hits_aug.module_minhu + hits_aug.module_maxhu) / 2) * hits_aug.value</p>\n  \n  <p>hits_aug[\"v\"] = ((hits_aug.ch1 + 0.5) * hits_aug.pitch_v - hits_aug.module_hv) * hits_aug.value</p>\n  \n  <p>group = hits_aug[[\"hit_id\", \"u\", \"v\", \"value\"]].groupby(\"hit_id\",as_index = False).sum()</p>\n  \n  <p>hits_uv = hits_aug.merge(group, on=\"hit_id\", suffixes=('','_sum'))</p>\n  \n  <p>hits_uv[\"est_x\"] = hits_uv.u_sum * hits_uv.rot_xu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_xv / hits_uv.value_sum + hits_uv.cx</p>\n  \n  <p>hits_uv[\"est_y\"] = hits_uv.u_sum * hits_uv.rot_yu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_yv / hits_uv.value_sum + hits_uv.cy</p>\n  \n  <p>hits_uv[\"est_z\"] = hits_uv.u_sum * hits_uv.rot_zu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_zv / \n  hits_uv.value_sum + hits_uv.cz</p>\n  \n  <p>result = hits_uv[[\"x\", \"y\", \"z\", \"est_x\", \"est_y\", \"est_z\"]].drop_duplicates()</p>\n</blockquote>",
      "rawMarkdown": "Here is how to properly calculate it:\n\n&gt;hits, cells, particles, truth = load_event('train_100_events/event000001000')\n\n&gt;detectors = pd.read_csv(\"detectors.csv\")\n\n&gt;hits_aug = hits.merge(detectors, on=[\"volume_id\", \"layer_id\", \"module_id\"]).merge(cells, on=\"hit_id\")\n\n&gt;hits_aug[\"u\"] = ((hits_aug.ch0 + 0.5) * hits_aug.pitch_u - (hits_aug.module_minhu + hits_aug.module_maxhu) / 2) * hits_aug.value\n\n&gt;hits_aug[\"v\"] = ((hits_aug.ch1 + 0.5) * hits_aug.pitch_v - hits_aug.module_hv) * hits_aug.value\n\n&gt;group = hits_aug[[\"hit_id\", \"u\", \"v\", \"value\"]].groupby(\"hit_id\",as_index = False).sum()\n\n&gt;hits_uv = hits_aug.merge(group, on=\"hit_id\", suffixes=('','_sum'))\n\n&gt;hits_uv[\"est_x\"] = hits_uv.u_sum * hits_uv.rot_xu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_xv / hits_uv.value_sum + hits_uv.cx\n\n&gt;hits_uv[\"est_y\"] = hits_uv.u_sum * hits_uv.rot_yu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_yv / hits_uv.value_sum + hits_uv.cy\n\n&gt;hits_uv[\"est_z\"] = hits_uv.u_sum * hits_uv.rot_zu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_zv / \nhits_uv.value_sum + hits_uv.cz\n\n&gt;result = hits_uv[[\"x\", \"y\", \"z\", \"est_x\", \"est_y\", \"est_z\"]].drop_duplicates()",
      "votes": null
    },
    {
      "id": "328837",
      "postDate": "05/15/2018 07:00:11",
      "content": "<p>Doesn't seem like it: with your code I get differences up to 5mm...</p>",
      "rawMarkdown": "Doesn't seem like it: with your code I get differences up to 5mm...",
      "votes": null
    },
    {
      "id": "328838",
      "postDate": "05/15/2018 07:03:06",
      "content": "<p>I also tried swapping averaging and rotation, but that still gives me same result as in <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/56684#328682\">this comment</a>...</p>",
      "rawMarkdown": "I also tried swapping averaging and rotation, but that still gives me same result as in [this comment][1]...\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/56684#328682",
      "votes": null
    },
    {
      "id": "328851",
      "postDate": "05/15/2018 07:26:49",
      "content": "<p>It's because of trapezoid ...</p>\n\n<p>result = hits_uv.loc[(hits_uv.module_maxhu == hits_uv.module_minhu),[\"hit_id\",\"x\", \"y\", \"z\", \"est_x\", \"est_y\", \"est_z\"]].drop_duplicates()</p>\n\n<p>(((result.x.round(4) - result.est_x.round(4))**2 + (result.y.round(4) - result.est_y.round(4))**2 + (result.z.round(4) - result.est_z.round(4))**2 + (result.z.round(4) - result.est_z.round(4))*<em>2) *</em> 0.5).mean()</p>\n\n<p>So the mean \"error\" in position for non-trapezoid is actually 1.6138009026212142e-2 which is probably generated by floating point calculations.</p>\n\n<p>If you look-up hit_id where the error is generated you'll see it comes from trapezoid and from long strip cells.</p>",
      "rawMarkdown": "It's because of trapezoid ...\n\nresult = hits_uv.loc[(hits_uv.module_maxhu == hits_uv.module_minhu),[\"hit_id\",\"x\", \"y\", \"z\", \"est_x\", \"est_y\", \"est_z\"]].drop_duplicates()\n\n(((result.x.round(4) - result.est_x.round(4))**2 + (result.y.round(4) - result.est_y.round(4))**2 + (result.z.round(4) - result.est_z.round(4))**2 + (result.z.round(4) - result.est_z.round(4))**2) ** 0.5).mean()\n\nSo the mean \"error\" in position for non-trapezoid is actually 1.6138009026212142e-2 which is probably generated by floating point calculations.\n\nIf you look-up hit_id where the error is generated you'll see it comes from trapezoid and from long strip cells.",
      "votes": null
    },
    {
      "id": "328852",
      "postDate": "05/15/2018 07:30:17",
      "content": "<p>Sorry but markdown is messing my code :-) you know the formula for distance ... mean it and you'll see the error is small.</p>",
      "rawMarkdown": "Sorry but markdown is messing my code :-) you know the formula for distance ... mean it and you'll see the error is small.",
      "votes": null
    },
    {
      "id": "328853",
      "postDate": "05/15/2018 07:36:55",
      "content": "<p>That's still same order of error as what I get, while my code works well for trapezoid. Actually, for some reason I get bigger errors for non-trapezoid cells than for trapezoid ones...</p>",
      "rawMarkdown": "That's still same order of error as what I get, while my code works well for trapezoid. Actually, for some reason I get bigger errors for non-trapezoid cells than for trapezoid ones...",
      "votes": null
    },
    {
      "id": "328854",
      "postDate": "05/15/2018 07:37:09",
      "content": "<p>What I find confusing is the difference between detectors.csv pitch values and the ones described in pdf. In PDF it states that the cells are square for pixel detectors (volume 8, 13, 17) and have a size of 0.05 by 0.05 yet the detectors file has values of 0.05 by 0.05625. For long strips it states 0.12 by 10.8 but in file it's 0.12 by 10.4. So that can be the cause of error.</p>",
      "rawMarkdown": "What I find confusing is the difference between detectors.csv pitch values and the ones described in pdf. In PDF it states that the cells are square for pixel detectors (volume 8, 13, 17) and have a size of 0.05 by 0.05 yet the detectors file has values of 0.05 by 0.05625. For long strips it states 0.12 by 10.8 but in file it's 0.12 by 10.4. So that can be the cause of error.",
      "votes": null
    },
    {
      "id": "328862",
      "postDate": "05/15/2018 07:43:35",
      "content": "<p>I would also mention that I think that at least some rot_xu and rot_yu are \"wrong\" as I've tried recreating the 3D model of detectors and I have some issues with modules with X or Y around 0 as they seem to be \"flipped\". </p>",
      "rawMarkdown": "I would also mention that I think that at least some rot_xu and rot_yu are \"wrong\" as I've tried recreating the 3D model of detectors and I have some issues with modules with X or Y around 0 as they seem to be \"flipped\".",
      "votes": null
    },
    {
      "id": "328875",
      "postDate": "05/15/2018 08:22:12",
      "content": "<p>Just checked the rotations (see updated kernel from the original post) - the orthogonality is fine. Also checked where the norm to each module is pointed to wrt vectors (cx, cy, 0) and (0, 0, cz) - the norm is always orthogonal to one of these.</p>",
      "rawMarkdown": "Just checked the rotations (see updated kernel from the original post) - the orthogonality is fine. Also checked where the norm to each module is pointed to wrt vectors (cx, cy, 0) and (0, 0, cz) - the norm is always orthogonal to one of these.",
      "votes": null
    },
    {
      "id": "329328",
      "postDate": "05/16/2018 08:39:29",
      "content": "<p>Hi, please take the values from the <code>detector.csv</code> file, they are properly calculated when the detector is written out - that's the problem with paper-style documentation:  the moment you write it, it becomes outdated. \nSorry for the inconvenience. </p>",
      "rawMarkdown": "Hi, please take the values from the `detector.csv` file, they are properly calculated when the detector is written out - that's the problem with paper-style documentation:  the moment you write it, it becomes outdated. \nSorry for the inconvenience.",
      "votes": null
    },
    {
      "id": "329333",
      "postDate": "05/16/2018 08:57:24",
      "content": "<p>Thank you Andreas, your <a href=\"https://www.kaggle.com/asalzburger/pixel-detector-cells\">Pixel Detector: Cells</a> kernel is excelent to explore Cells. @SiLiKhon have a look at it ;-)</p>",
      "rawMarkdown": "Thank you Andreas, your [Pixel Detector: Cells][1] kernel is excelent to explore Cells. @SiLiKhon have a look at it ;-)\n\n\n  [1]: https://www.kaggle.com/asalzburger/pixel-detector-cells",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 328677,
      "author_name": "profetul",
      "author_url": "",
      "post_date": "05/14/2018 21:58:50",
      "content": "<p><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf\">https://storage.googleapis.com/kaggle-forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf</a>.  See Appendix A. I would suggest trying a code where you calculate the estimated (X, Y) by doing first (m0,m1) and apply rot after that - not before.  </p>\n\n<p>I've also noticed you're ignoring rot_xw, rot_yw and rot_zw - the cell doesn't have a thickness of 0! It's true that the \"ch3\" is not provide, but maybe the W position is -module_T? So W can alter the value of X, Y, Z. </p>\n\n<p>In theory the distance should be then 0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 328682,
          "author_name": "silikhon",
          "author_url": "",
          "post_date": "05/14/2018 22:36:31",
          "content": "<p>Thanks for the suggestion.</p>\n\n<p>I doubt that order of rotation and averaging matters here: averaging is a linear operation and it commutes with rotation - if only it has some effects of finite numeric precision, I'll try and check that later.\nWhat does matter though is that I average the <em>distance</em> between cell centers and provided hit positions, rather than make an average of dx, dy and dz separately. I've tried this out and got the following dist:\n<img src=\"https://cernbox.cern.ch/index.php/s/MOl03nMQ8fr1XO4/download\" alt=\"mean dist\"></p>\n\n<p>I'll try swapping the rotation and averaging to see if it goes to 0 exactly.\nNote that smallest module width is 0.150 which is much bigger then the discrepancy I'm having here, so it can't be it.</p>\n\n<p>Thanks for pointing me to the right spot!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328789,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/15/2018 04:49:14",
          "content": "<p>I think you're wrong about distances. Something else I've noticed in your code is that you don't take into account the fact that some modules are trapezoid and assume the number of cells on the smaller base to be the same as on the bigger base and it might not be the case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328799,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/15/2018 05:09:15",
          "content": "<p>Actually I still can't get my head around distances. You are saying that the sum of distance between a point and other points is the same as distance of that point to the point resulted from the sum of points. I don't think that's true. Let's take 0 as our point and -1 and 1 the other points. Distance from 0 to -1 is 1, to 1 is 1, sum is 2. Distance to sum of points (-1 + 1) is 0. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328816,
          "author_name": "silikhon",
          "author_url": "",
          "post_date": "05/15/2018 06:11:14",
          "content": "<p>Regarding the trapezoid shape, I'm following this comment <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/56266#327664\">https://www.kaggle.com/c/trackml-particle-identification/discussion/56266#327664</a> by Andreas, stating that grid is still rectangular for simplicity.\nI'm assuming this means the cells are labeled as if they were in a rectangular module of shape (2 * module_hv,  2 * module_maxhu), and just indexes of cells cut off by the trapezoidal corners never occur. That actually seems to be true. Here's the same plot as before just for the trapezoidal modules:\n<img src=\"https://cernbox.cern.ch/index.php/s/SweosZKPJozLkuQ/download\" alt=\"mean dist for trapezoid\">\nNote that smallest possible step between cells in trapezoid modules is 0.08mm, which is much bigger than the diff I get on this hist. If I missed cell position by one or more steps, you'd see it in this figure as an entry of 0.08 or greater.</p>\n\n<p>Now on your second comment:</p>\n\n<pre><code>Let's take 0 as our point and -1 and 1 the other points. Distance from 0 to -1 is 1, to 1 is 1, sum is 2. Distance to sum of points (-1 + 1) is 0.\n</code></pre>\n\n<p>this was exactly the problem with the code in the kernel. As I said before, once you change that to average of dx, dy and dz separately, you get the distribution as in my last comment. Applying that to your example: dx between -1 and 0 is -1, between 1 and 0 is 1 which makes sum(dx) = -1 + 1 = 0, which is same as dx between (-1 + 1) and 0.</p>\n\n<p>P.S.\nI haven't updated the kernel yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328824,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/15/2018 06:31:59",
          "content": "<p>Here is how to properly calculate it:</p>\n\n<blockquote>\n  <p>hits, cells, particles, truth = load_event('train_100_events/event000001000')</p>\n  \n  <p>detectors = pd.read_csv(\"detectors.csv\")</p>\n  \n  <p>hits_aug = hits.merge(detectors, on=[\"volume_id\", \"layer_id\", \"module_id\"]).merge(cells, on=\"hit_id\")</p>\n  \n  <p>hits_aug[\"u\"] = ((hits_aug.ch0 + 0.5) * hits_aug.pitch_u - (hits_aug.module_minhu + hits_aug.module_maxhu) / 2) * hits_aug.value</p>\n  \n  <p>hits_aug[\"v\"] = ((hits_aug.ch1 + 0.5) * hits_aug.pitch_v - hits_aug.module_hv) * hits_aug.value</p>\n  \n  <p>group = hits_aug[[\"hit_id\", \"u\", \"v\", \"value\"]].groupby(\"hit_id\",as_index = False).sum()</p>\n  \n  <p>hits_uv = hits_aug.merge(group, on=\"hit_id\", suffixes=('','_sum'))</p>\n  \n  <p>hits_uv[\"est_x\"] = hits_uv.u_sum * hits_uv.rot_xu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_xv / hits_uv.value_sum + hits_uv.cx</p>\n  \n  <p>hits_uv[\"est_y\"] = hits_uv.u_sum * hits_uv.rot_yu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_yv / hits_uv.value_sum + hits_uv.cy</p>\n  \n  <p>hits_uv[\"est_z\"] = hits_uv.u_sum * hits_uv.rot_zu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_zv / \n  hits_uv.value_sum + hits_uv.cz</p>\n  \n  <p>result = hits_uv[[\"x\", \"y\", \"z\", \"est_x\", \"est_y\", \"est_z\"]].drop_duplicates()</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328837,
          "author_name": "silikhon",
          "author_url": "",
          "post_date": "05/15/2018 07:00:11",
          "content": "<p>Doesn't seem like it: with your code I get differences up to 5mm...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328838,
          "author_name": "silikhon",
          "author_url": "",
          "post_date": "05/15/2018 07:03:06",
          "content": "<p>I also tried swapping averaging and rotation, but that still gives me same result as in <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/56684#328682\">this comment</a>...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328851,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/15/2018 07:26:49",
          "content": "<p>It's because of trapezoid ...</p>\n\n<p>result = hits_uv.loc[(hits_uv.module_maxhu == hits_uv.module_minhu),[\"hit_id\",\"x\", \"y\", \"z\", \"est_x\", \"est_y\", \"est_z\"]].drop_duplicates()</p>\n\n<p>(((result.x.round(4) - result.est_x.round(4))**2 + (result.y.round(4) - result.est_y.round(4))**2 + (result.z.round(4) - result.est_z.round(4))**2 + (result.z.round(4) - result.est_z.round(4))*<em>2) *</em> 0.5).mean()</p>\n\n<p>So the mean \"error\" in position for non-trapezoid is actually 1.6138009026212142e-2 which is probably generated by floating point calculations.</p>\n\n<p>If you look-up hit_id where the error is generated you'll see it comes from trapezoid and from long strip cells.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328852,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/15/2018 07:30:17",
          "content": "<p>Sorry but markdown is messing my code :-) you know the formula for distance ... mean it and you'll see the error is small.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328853,
          "author_name": "silikhon",
          "author_url": "",
          "post_date": "05/15/2018 07:36:55",
          "content": "<p>That's still same order of error as what I get, while my code works well for trapezoid. Actually, for some reason I get bigger errors for non-trapezoid cells than for trapezoid ones...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328854,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/15/2018 07:37:09",
          "content": "<p>What I find confusing is the difference between detectors.csv pitch values and the ones described in pdf. In PDF it states that the cells are square for pixel detectors (volume 8, 13, 17) and have a size of 0.05 by 0.05 yet the detectors file has values of 0.05 by 0.05625. For long strips it states 0.12 by 10.8 but in file it's 0.12 by 10.4. So that can be the cause of error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328862,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/15/2018 07:43:35",
          "content": "<p>I would also mention that I think that at least some rot_xu and rot_yu are \"wrong\" as I've tried recreating the 3D model of detectors and I have some issues with modules with X or Y around 0 as they seem to be \"flipped\". </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 328875,
          "author_name": "silikhon",
          "author_url": "",
          "post_date": "05/15/2018 08:22:12",
          "content": "<p>Just checked the rotations (see updated kernel from the original post) - the orthogonality is fine. Also checked where the norm to each module is pointed to wrt vectors (cx, cy, 0) and (0, 0, cz) - the norm is always orthogonal to one of these.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 329328,
          "author_name": "asalzburger",
          "author_url": "",
          "post_date": "05/16/2018 08:39:29",
          "content": "<p>Hi, please take the values from the <code>detector.csv</code> file, they are properly calculated when the detector is written out - that's the problem with paper-style documentation:  the moment you write it, it becomes outdated. \nSorry for the inconvenience. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 329333,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/16/2018 08:57:24",
          "content": "<p>Thank you Andreas, your <a href=\"https://www.kaggle.com/asalzburger/pixel-detector-cells\">Pixel Detector: Cells</a> kernel is excelent to explore Cells. @SiLiKhon have a look at it ;-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "328129": "Dear organizers,\n\nFirstly, thanks a lot for making such a competition possible!\n\nNow, I'm trying to get the xyz coordinates of the cells using ch0 and ch1 parameters.\n[Here's the code I wrote for that][1]. What bugs me is that the mean difference between the provided hit xyz coordinates and the calculated coords of the cells involved (weighted by the cell value, assuming it's proportional to the energy deposition within that cell) has some weird distribution:\n![Weighted mean distance between the calculated cell center and hit position][2]\n(This is a single-event distribution, evt # 000001100 in particular - but I've tried various events and the distribution looks similarly)\n\nDoes this simply mean that hit position is defined in a more complicated way than just cell position mean, or is there a problem in my cell position calculation?\n\nWould appreciate any help. Thanks!\n\n  [1]: https://www.kaggle.com/silikhon/cellposition\n  [2]: https://cernbox.cern.ch/index.php/s/A6cDax9t1nMOyyU/download",
    "328677": "https://kaggle2.blob.core.windows.net/forum-message-attachments/321278/9331/trackml-participant-document-particle-v1.0.pdf.  See Appendix A. I would suggest trying a code where you calculate the estimated (X, Y) by doing first (m0,m1) and apply rot after that - not before.  \n\nI've also noticed you're ignoring rot_xw, rot_yw and rot_zw - the cell doesn't have a thickness of 0! It's true that the \"ch3\" is not provide, but maybe the W position is -module_T? So W can alter the value of X, Y, Z. \n\nIn theory the distance should be then 0.",
    "328682": "Thanks for the suggestion.\n\nI doubt that order of rotation and averaging matters here: averaging is a linear operation and it commutes with rotation - if only it has some effects of finite numeric precision, I'll try and check that later.\nWhat does matter though is that I average the *distance* between cell centers and provided hit positions, rather than make an average of dx, dy and dz separately. I've tried this out and got the following dist:\n![mean dist][1]\n\nI'll try swapping the rotation and averaging to see if it goes to 0 exactly.\nNote that smallest module width is 0.150 which is much bigger then the discrepancy I'm having here, so it can't be it.\n\nThanks for pointing me to the right spot!\n\n  [1]: https://cernbox.cern.ch/index.php/s/MOl03nMQ8fr1XO4/download",
    "328789": "I think you're wrong about distances. Something else I've noticed in your code is that you don't take into account the fact that some modules are trapezoid and assume the number of cells on the smaller base to be the same as on the bigger base and it might not be the case.",
    "328799": "Actually I still can't get my head around distances. You are saying that the sum of distance between a point and other points is the same as distance of that point to the point resulted from the sum of points. I don't think that's true. Let's take 0 as our point and -1 and 1 the other points. Distance from 0 to -1 is 1, to 1 is 1, sum is 2. Distance to sum of points (-1 + 1) is 0.",
    "328816": "Regarding the trapezoid shape, I'm following this comment https://www.kaggle.com/c/trackml-particle-identification/discussion/56266#327664 by Andreas, stating that grid is still rectangular for simplicity.\nI'm assuming this means the cells are labeled as if they were in a rectangular module of shape (2 * module_hv,  2 * module_maxhu), and just indexes of cells cut off by the trapezoidal corners never occur. That actually seems to be true. Here's the same plot as before just for the trapezoidal modules:\n![mean dist for trapezoid][1]\nNote that smallest possible step between cells in trapezoid modules is 0.08mm, which is much bigger than the diff I get on this hist. If I missed cell position by one or more steps, you'd see it in this figure as an entry of 0.08 or greater.\n\nNow on your second comment:\n\n    Let's take 0 as our point and -1 and 1 the other points. Distance from 0 to -1 is 1, to 1 is 1, sum is 2. Distance to sum of points (-1 + 1) is 0.\n\nthis was exactly the problem with the code in the kernel. As I said before, once you change that to average of dx, dy and dz separately, you get the distribution as in my last comment. Applying that to your example: dx between -1 and 0 is -1, between 1 and 0 is 1 which makes sum(dx) = -1 + 1 = 0, which is same as dx between (-1 + 1) and 0.\n\nP.S.\nI haven't updated the kernel yet.\n\n  [1]: https://cernbox.cern.ch/index.php/s/SweosZKPJozLkuQ/download",
    "328824": "Here is how to properly calculate it:\n\n&gt;hits, cells, particles, truth = load_event('train_100_events/event000001000')\n\n&gt;detectors = pd.read_csv(\"detectors.csv\")\n\n&gt;hits_aug = hits.merge(detectors, on=[\"volume_id\", \"layer_id\", \"module_id\"]).merge(cells, on=\"hit_id\")\n\n&gt;hits_aug[\"u\"] = ((hits_aug.ch0 + 0.5) * hits_aug.pitch_u - (hits_aug.module_minhu + hits_aug.module_maxhu) / 2) * hits_aug.value\n\n&gt;hits_aug[\"v\"] = ((hits_aug.ch1 + 0.5) * hits_aug.pitch_v - hits_aug.module_hv) * hits_aug.value\n\n&gt;group = hits_aug[[\"hit_id\", \"u\", \"v\", \"value\"]].groupby(\"hit_id\",as_index = False).sum()\n\n&gt;hits_uv = hits_aug.merge(group, on=\"hit_id\", suffixes=('','_sum'))\n\n&gt;hits_uv[\"est_x\"] = hits_uv.u_sum * hits_uv.rot_xu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_xv / hits_uv.value_sum + hits_uv.cx\n\n&gt;hits_uv[\"est_y\"] = hits_uv.u_sum * hits_uv.rot_yu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_yv / hits_uv.value_sum + hits_uv.cy\n\n&gt;hits_uv[\"est_z\"] = hits_uv.u_sum * hits_uv.rot_zu / hits_uv.value_sum + hits_uv.v_sum * hits_uv.rot_zv / \nhits_uv.value_sum + hits_uv.cz\n\n&gt;result = hits_uv[[\"x\", \"y\", \"z\", \"est_x\", \"est_y\", \"est_z\"]].drop_duplicates()",
    "328837": "Doesn't seem like it: with your code I get differences up to 5mm...",
    "328838": "I also tried swapping averaging and rotation, but that still gives me same result as in [this comment][1]...\n\n  [1]: https://www.kaggle.com/c/trackml-particle-identification/discussion/56684#328682",
    "328851": "It's because of trapezoid ...\n\nresult = hits_uv.loc[(hits_uv.module_maxhu == hits_uv.module_minhu),[\"hit_id\",\"x\", \"y\", \"z\", \"est_x\", \"est_y\", \"est_z\"]].drop_duplicates()\n\n(((result.x.round(4) - result.est_x.round(4))**2 + (result.y.round(4) - result.est_y.round(4))**2 + (result.z.round(4) - result.est_z.round(4))**2 + (result.z.round(4) - result.est_z.round(4))**2) ** 0.5).mean()\n\nSo the mean \"error\" in position for non-trapezoid is actually 1.6138009026212142e-2 which is probably generated by floating point calculations.\n\nIf you look-up hit_id where the error is generated you'll see it comes from trapezoid and from long strip cells.",
    "328852": "Sorry but markdown is messing my code :-) you know the formula for distance ... mean it and you'll see the error is small.",
    "328853": "That's still same order of error as what I get, while my code works well for trapezoid. Actually, for some reason I get bigger errors for non-trapezoid cells than for trapezoid ones...",
    "328854": "What I find confusing is the difference between detectors.csv pitch values and the ones described in pdf. In PDF it states that the cells are square for pixel detectors (volume 8, 13, 17) and have a size of 0.05 by 0.05 yet the detectors file has values of 0.05 by 0.05625. For long strips it states 0.12 by 10.8 but in file it's 0.12 by 10.4. So that can be the cause of error.",
    "328862": "I would also mention that I think that at least some rot_xu and rot_yu are \"wrong\" as I've tried recreating the 3D model of detectors and I have some issues with modules with X or Y around 0 as they seem to be \"flipped\".",
    "328875": "Just checked the rotations (see updated kernel from the original post) - the orthogonality is fine. Also checked where the norm to each module is pointed to wrt vectors (cx, cy, 0) and (0, 0, cz) - the norm is always orthogonal to one of these.",
    "329328": "Hi, please take the values from the `detector.csv` file, they are properly calculated when the detector is written out - that's the problem with paper-style documentation:  the moment you write it, it becomes outdated. \nSorry for the inconvenience.",
    "329333": "Thank you Andreas, your [Pixel Detector: Cells][1] kernel is excelent to explore Cells. @SiLiKhon have a look at it ;-)\n\n\n  [1]: https://www.kaggle.com/asalzburger/pixel-detector-cells"
  },
  "source": "meta"
}