{
  "id": 60824,
  "title": "Can CERN employees participate?",
  "url": "/competitions/trackml-particle-identification/discussion/60824",
  "author_name": "",
  "post_date": "2018-07-10T11:22:06.459981Z",
  "votes": 2,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Recent comments made me wonder about CERN members joining late and crushing the LB.  In most, if not all Kaggle competitions I entered, people working for the sponsor or for an affiliated entity of the sponsor cannot enter the competition or if they can, they are not eligible to prizes.  The relevant rule is 2.B but it is not clear at all to me.   Seems associated members are allowed to enter and claim prize.  I hope they don't have access to the test data and/or the code of the simulator used to generate the events.  </p>\n\n<p>And if they can enter, then why aren't we seeing any?  I also hope they don' t wait for the last minute to submit.</p>",
  "messages": [
    {
      "id": "354861",
      "postDate": "07/10/2018 11:22:06",
      "content": "<p>Recent comments made me wonder about CERN members joining late and crushing the LB.  In most, if not all Kaggle competitions I entered, people working for the sponsor or for an affiliated entity of the sponsor cannot enter the competition or if they can, they are not eligible to prizes.  The relevant rule is 2.B but it is not clear at all to me.   Seems associated members are allowed to enter and claim prize.  I hope they don't have access to the test data and/or the code of the simulator used to generate the events.  </p>\n\n<p>And if they can enter, then why aren't we seeing any?  I also hope they don' t wait for the last minute to submit.</p>",
      "rawMarkdown": "Recent comments made me wonder about CERN members joining late and crushing the LB.  In most, if not all Kaggle competitions I entered, people working for the sponsor or for an affiliated entity of the sponsor cannot enter the competition or if they can, they are not eligible to prizes.  The relevant rule is 2.B but it is not clear at all to me.   Seems associated members are allowed to enter and claim prize.  I hope they don't have access to the test data and/or the code of the simulator used to generate the events.  \n\nAnd if they can enter, then why aren't we seeing any?  I also hope they don' t wait for the last minute to submit.",
      "votes": null
    },
    {
      "id": "354939",
      "postDate": "07/10/2018 14:23:14",
      "content": "<p>Hi, CERN employees can participate but cannot claim any price. CERN being the largest particle physics laboratory in the world, more than half of all professional particle physicists have the so-called \"CERN user\" status, these people are not excluded. This would be equivalent for e.g. a DNA labelling competition to exclude most people with a molecular biology PhD.    However, for sure, these people (nor CERN employees in fact) have had no privileged access to the master simulation code  nor the test data. The only people to have access to the simulator are the organizers. And to be extra safe, only two of them have access to the test data truth.</p>",
      "rawMarkdown": "Hi, CERN employees can participate but cannot claim any price. CERN being the largest particle physics laboratory in the world, more than half of all professional particle physicists have the so-called \"CERN user\" status, these people are not excluded. This would be equivalent for e.g. a DNA labelling competition to exclude most people with a molecular biology PhD.    However, for sure, these people (nor CERN employees in fact) have had no privileged access to the master simulation code  nor the test data. The only people to have access to the simulator are the organizers. And to be extra safe, only two of them have access to the test data truth.",
      "votes": null
    },
    {
      "id": "354956",
      "postDate": "07/10/2018 14:55:01",
      "content": "<p>Thank you, makes perfect sense.</p>",
      "rawMarkdown": "Thank you, makes perfect sense.",
      "votes": null
    },
    {
      "id": "355056",
      "postDate": "07/10/2018 20:46:11",
      "content": "<blockquote>\n  <p>And if they can enter, then why aren't we seeing any? I also hope they don' t wait for the last minute to submit.</p>\n</blockquote>\n\n<p>One can make some hypotheses, mostly pessimistic ones: Maybe the professionals do not think there is any original approach which has a chance against the established ones. In that case a win would not be academically or scientifically valuable.</p>",
      "rawMarkdown": "&gt; And if they can enter, then why aren't we seeing any? I also hope they don' t wait for the last minute to submit.\n\nOne can make some hypotheses, mostly pessimistic ones: Maybe the professionals do not think there is any original approach which has a chance against the established ones. In that case a win would not be academically or scientifically valuable.",
      "votes": null
    },
    {
      "id": "355068",
      "postDate": "07/10/2018 21:29:34",
      "content": "<p>A positive spin is that the professionals do not think an original approach coming from their own community is likely to emerge in a few months. Professionals also know that developing software for a new detector (like the one we simulated for the challenge, which was designed specially for the challenge, nobody has seen it before, even though it mimics broadly ATLAS and CMS future tracker ) takes years rather than months. For example we are all stunned that DBSCAN which we have never ever seen used in tracking, which we provided just because it was an almost one-liner, could be carried and improved to such significant score. We would have bet more on Hough transform which has been completely ignored, as far as we can tell from the forum. We're very impressed by the performances reached and we're looking forward to have detailed on the algorithm used.</p>",
      "rawMarkdown": "A positive spin is that the professionals do not think an original approach coming from their own community is likely to emerge in a few months. Professionals also know that developing software for a new detector (like the one we simulated for the challenge, which was designed specially for the challenge, nobody has seen it before, even though it mimics broadly ATLAS and CMS future tracker ) takes years rather than months. For example we are all stunned that DBSCAN which we have never ever seen used in tracking, which we provided just because it was an almost one-liner, could be carried and improved to such significant score. We would have bet more on Hough transform which has been completely ignored, as far as we can tell from the forum. We're very impressed by the performances reached and we're looking forward to have detailed on the algorithm used.",
      "votes": null
    },
    {
      "id": "355193",
      "postDate": "07/11/2018 06:59:43",
      "content": "<p>I don't think hough transform is ignored here.  It is implemented via 'helix unrolling', but it amounts to the same AFAIK.</p>\n\n<blockquote>\n  <p>it mimics broadly ATLAS and CMS future tracker</p>\n</blockquote>\n\n<p>Thanks for confirming that this data is different from what researchers in the field have been seeing so far.  It was said in some of the shared material, but people did not notice it really.  This gives some hope for participants I guess.  </p>",
      "rawMarkdown": "I don't think hough transform is ignored here.  It is implemented via 'helix unrolling', but it amounts to the same AFAIK.\n\n&gt; it mimics broadly ATLAS and CMS future tracker\n\nThanks for confirming that this data is different from what researchers in the field have been seeing so far.  It was said in some of the shared material, but people did not notice it really.  This gives some hope for participants I guess.",
      "votes": null
    },
    {
      "id": "355210",
      "postDate": "07/11/2018 07:52:46",
      "content": "<blockquote>\n  <p>I don't think hough transform is ignored here. It is implemented via 'helix unrolling',</p>\n</blockquote>\n\n<p>I guess what David is referring to is the histogramming technique, not only the transform. My take on it would be that histogramming approaches are ignored for good reasons, but I may be wrong.</p>\n\n<p>As you imply, probably every single serious competitor is doing <em>some</em> kind of coordinate transform in the broadest sense to make the problem approachable. One might then say that massive helix unrolling combined with a merging strategy is a kind of hidden histogram approach, depending on how the merging works exactly.</p>",
      "rawMarkdown": "&gt; I don't think hough transform is ignored here. It is implemented via 'helix unrolling',\n\nI guess what David is referring to is the histogramming technique, not only the transform. My take on it would be that histogramming approaches are ignored for good reasons, but I may be wrong.\n\nAs you imply, probably every single serious competitor is doing *some* kind of coordinate transform in the broadest sense to make the problem approachable. One might then say that massive helix unrolling combined with a merging strategy is a kind of hidden histogram approach, depending on how the merging works exactly.",
      "votes": null
    },
    {
      "id": "355214",
      "postDate": "07/11/2018 08:03:02",
      "content": "<p>I view histogramming as one way to do clustering ;)  The point is that using some polar coordinates and transforming them is used by many here, and this is the general idea of hough transform.  I agree that the specific way hough transform is used in one of the shared kernels may not be reused much.</p>",
      "rawMarkdown": "I view histogramming as one way to do clustering ;)  The point is that using some polar coordinates and transforming them is used by many here, and this is the general idea of hough transform.  I agree that the specific way hough transform is used in one of the shared kernels may not be reused much.",
      "votes": null
    },
    {
      "id": "355252",
      "postDate": "07/11/2018 09:45:37",
      "content": "<p>It also seems to me that \"dbscan &amp; helix-unrolling\" is mainly equivalent to Hough transform. I preferred Hough transform at the beginning, because it seems intuitively right. But after quite some work, I turned to the \"helix-unrolling\"-code, which worked better, and seems to have the same substance:</p>\n\n<p>Dbscan takes the role of putting the hits into bins (as done via the \"ComboDigi\" column at the Hough transform code).</p>\n\n<p>The \"helix-unrolling\" seems equivalent to what Hough transform is really about. Each iteration at \"helix-unrolling\" does a particular transformation to the hits. Each iteration yields a particular set of long tracks: tracks which have a particular speed along the z-axis (via \"dz\"), or (e.g. if another specific main feature is used) whose mid-points, when projected onto the x-y plane, have a particular angle. I think the Hough transform code does the latter. </p>\n\n<p>Those methods, as well as the main feature (which is adjusted via the iteration variable), which is used, seem to be equivalent, but hyper-parameter optimization and computational speed might be (much) better for one method (or main feature) vs another.</p>",
      "rawMarkdown": "It also seems to me that \"dbscan &amp; helix-unrolling\" is mainly equivalent to Hough transform. I preferred Hough transform at the beginning, because it seems intuitively right. But after quite some work, I turned to the \"helix-unrolling\"-code, which worked better, and seems to have the same substance:\n\nDbscan takes the role of putting the hits into bins (as done via the \"ComboDigi\" column at the Hough transform code).\n\nThe \"helix-unrolling\" seems equivalent to what Hough transform is really about. Each iteration at \"helix-unrolling\" does a particular transformation to the hits. Each iteration yields a particular set of long tracks: tracks which have a particular speed along the z-axis (via \"dz\"), or (e.g. if another specific main feature is used) whose mid-points, when projected onto the x-y plane, have a particular angle. I think the Hough transform code does the latter. \n\nThose methods, as well as the main feature (which is adjusted via the iteration variable), which is used, seem to be equivalent, but hyper-parameter optimization and computational speed might be (much) better for one method (or main feature) vs another.",
      "votes": null
    },
    {
      "id": "355282",
      "postDate": "07/11/2018 11:29:29",
      "content": "<p>@Edwin, found this blog post about this competition, so I guess you're a German speaker (too). \n<a href=\"https://motls.blogspot.com/2018/07/our-edwin-steiner-is-current-leader-in.html\">https://motls.blogspot.com/2018/07/our-edwin-steiner-is-current-leader-in.html</a></p>\n\n<pre><code>Even if Outrunner improves his or her score, it's not the end of the hopes because, as Edwin says, \"Es kommt der Tag\". Or, more precisely, \"Aber es ist noch nicht aller Tage Abend!\"\n</code></pre>",
      "rawMarkdown": "Edwin, found this blog post about this competition, so I guess you're a German speaker (too). \nhttps://motls.blogspot.com/2018/07/our-edwin-steiner-is-current-leader-in.html\n\n    Even if Outrunner improves his or her score, it's not the end of the hopes because, as Edwin says, \"Es kommt der Tag\". Or, more precisely, \"Aber es ist noch nicht aller Tage Abend!\"",
      "votes": null
    },
    {
      "id": "355469",
      "postDate": "07/11/2018 18:36:55",
      "content": "<p>@CPMP,</p>\n\n<blockquote>\n  <p>I view histogramming as one way to do clustering</p>\n</blockquote>\n\n<p>Makes sense, but the clustering only covers two dimensions of the track parameter histogram, doesn't it? The other dimensions then need to be dealt with by ensembling and merging clustering results in such approaches, for example.</p>",
      "rawMarkdown": "CPMP,\n\n&gt; I view histogramming as one way to do clustering\n\nMakes sense, but the clustering only covers two dimensions of the track parameter histogram, doesn't it? The other dimensions then need to be dealt with by ensembling and merging clustering results in such approaches, for example.",
      "votes": null
    },
    {
      "id": "355572",
      "postDate": "07/12/2018 02:29:09",
      "content": "<blockquote>\n  <p>clustering only covers two dimensions of the track parameter histogram, doesn't it?</p>\n</blockquote>\n\n<p>I am not sure I get what you mean.  You can do clustering with as many dimensions you want.  I'm using 3 dimensions for instance so far.</p>",
      "rawMarkdown": "&gt; clustering only covers two dimensions of the track parameter histogram, doesn't it?\n\nI am not sure I get what you mean.  You can do clustering with as many dimensions you want.  I'm using 3 dimensions for instance so far.",
      "votes": null
    },
    {
      "id": "355723",
      "postDate": "07/12/2018 08:57:35",
      "content": "<p>The space in which you are clustering can have any dimension, that's true. I think I've seen an example kernel using 5 dimensions. However, as long as you are clustering hits and the clustering coordinates are smooth functions of (x,y,z), the points will lie on a 3-dimensional manifold in the higher-dimensional space. (What I'm saying here does not apply to clustering of more complicated derived objects, which has been suggested in at least on discussion thread.) Typically you will want to project one dimension away either completely or approximately because the separation of hits along the trajectory should not keep them from being clustered together. That's what the example kernels do: They effectively project the points onto a 2-dimensional manifold in e.g. 5-dimensional metric space. In the end, it is equivalent to a 2-dimensional clustering problem with a custom distance function. That's what I meant.</p>\n\n<p>All or some of what I said may not apply if you use clustering in a more sophisticated or indirect way. I'm referring to the approaches put forward in the public kernels.</p>\n\n<p>Note: What I call \"projection\" here may be highly non-linear functions. I do not mean linear projection only.</p>",
      "rawMarkdown": "The space in which you are clustering can have any dimension, that's true. I think I've seen an example kernel using 5 dimensions. However, as long as you are clustering hits and the clustering coordinates are smooth functions of (x,y,z), the points will lie on a 3-dimensional manifold in the higher-dimensional space. (What I'm saying here does not apply to clustering of more complicated derived objects, which has been suggested in at least on discussion thread.) Typically you will want to project one dimension away either completely or approximately because the separation of hits along the trajectory should not keep them from being clustered together. That's what the example kernels do: They effectively project the points onto a 2-dimensional manifold in e.g. 5-dimensional metric space. In the end, it is equivalent to a 2-dimensional clustering problem with a custom distance function. That's what I meant.\n\nAll or some of what I said may not apply if you use clustering in a more sophisticated or indirect way. I'm referring to the approaches put forward in the public kernels.\n\nNote: What I call \"projection\" here may be highly non-linear functions. I do not mean linear projection only.",
      "votes": null
    },
    {
      "id": "355792",
      "postDate": "07/12/2018 11:20:28",
      "content": "<p>Agreed.</p>",
      "rawMarkdown": "Agreed.",
      "votes": null
    },
    {
      "id": "355799",
      "postDate": "07/12/2018 11:32:19",
      "content": "<p>Hough transform is more than helix unrolling. As indicated in <a href=\"https://www.kaggle.com/mikhailhushchyn/hough-transform\">https://www.kaggle.com/mikhailhushchyn/hough-transform</a> the key thing is that each point becomes a curve in the image space (2 dimensions in the kernel, but could have as many dimension as there are degrees of freedom in the object to be found), then the clustering is done in the image space (looking for intersection of the curves), before going back to the normal space.  While I have the impression that the evolved DBSCAN techniques mentioned in the forum are always clustering in some transformation of the normal space (where a point is always a point).  I understand this is done multiple times with variant of the transformation, then tracks found are ensembled.   </p>",
      "rawMarkdown": "Hough transform is more than helix unrolling. As indicated in https://www.kaggle.com/mikhailhushchyn/hough-transform the key thing is that each point becomes a curve in the image space (2 dimensions in the kernel, but could have as many dimension as there are degrees of freedom in the object to be found), then the clustering is done in the image space (looking for intersection of the curves), before going back to the normal space.  While I have the impression that the evolved DBSCAN techniques mentioned in the forum are always clustering in some transformation of the normal space (where a point is always a point).  I understand this is done multiple times with variant of the transformation, then tracks found are ensembled.",
      "votes": null
    },
    {
      "id": "355811",
      "postDate": "07/12/2018 11:56:35",
      "content": "<p>You're right, my bad.   I'm actually using a similar approach in one step of my approach, i.e. moving hits to curves in some space and finding intersections.  I missed that Hough transform was similar.</p>",
      "rawMarkdown": "You're right, my bad.   I'm actually using a similar approach in one step of my approach, i.e. moving hits to curves in some space and finding intersections.  I missed that Hough transform was similar.",
      "votes": null
    },
    {
      "id": "355813",
      "postDate": "07/12/2018 11:58:27",
      "content": "<p>I think I do not agree that it is \"more\". In my understanding: At \"helix unrolling\" each hit also becomes a curve (the collection of hits you get by transforming this hit for every different angle). Then you also look for intersection of those curves (=for one specific angle, you find 10 different hits whose hit transformation for this angle lie in a small cluster).</p>\n\n<p>Anyway, this discussion is more because of intellectual curiosity I guess ;)</p>",
      "rawMarkdown": "I think I do not agree that it is \"more\". In my understanding: At \"helix unrolling\" each hit also becomes a curve (the collection of hits you get by transforming this hit for every different angle). Then you also look for intersection of those curves (=for one specific angle, you find 10 different hits whose hit transformation for this angle lie in a small cluster).\n\nAnyway, this discussion is more because of intellectual curiosity I guess ;)",
      "votes": null
    },
    {
      "id": "355818",
      "postDate": "07/12/2018 12:09:00",
      "content": "<p>OK as you@train2018  explain it.</p>",
      "rawMarkdown": "OK as you@train2018  explain it.",
      "votes": null
    },
    {
      "id": "356321",
      "postDate": "07/13/2018 10:39:26",
      "content": "<p>@CPMP  Your work was too supportive sir. I have learnt a lot of things from your kernels and Your topics Sir.   Can i be an Collaborator for your team. So that i can learn some things from you sir</p>",
      "rawMarkdown": "CPMP  Your work was too supportive sir. I have learnt a lot of things from your kernels and Your topics Sir.   Can i be an Collaborator for your team. So that i can learn some things from you sir",
      "votes": null
    },
    {
      "id": "356544",
      "postDate": "07/13/2018 20:48:10",
      "content": "<p>@dineshbarri_1997, thanks a lot for your kind words.  I think I'll try to see how far I can go without teaming here.  Unlike Grzegorz, I still hope that a non specialist can fare well here.</p>",
      "rawMarkdown": "dineshbarri_1997, thanks a lot for your kind words.  I think I'll try to see how far I can go without teaming here.  Unlike Grzegorz, I still hope that a non specialist can fare well here.",
      "votes": null
    },
    {
      "id": "356550",
      "postDate": "07/13/2018 20:52:39",
      "content": "<p>Thank you so muchh sir. . Its privileged talking with you sir </p>",
      "rawMarkdown": "Thank you so muchh sir. . Its privileged talking with you sir",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 354939,
      "author_name": "droussea",
      "author_url": "",
      "post_date": "07/10/2018 14:23:14",
      "content": "<p>Hi, CERN employees can participate but cannot claim any price. CERN being the largest particle physics laboratory in the world, more than half of all professional particle physicists have the so-called \"CERN user\" status, these people are not excluded. This would be equivalent for e.g. a DNA labelling competition to exclude most people with a molecular biology PhD.    However, for sure, these people (nor CERN employees in fact) have had no privileged access to the master simulation code  nor the test data. The only people to have access to the simulator are the organizers. And to be extra safe, only two of them have access to the test data truth.</p>",
      "votes": null,
      "replies": [
        {
          "id": 354956,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/10/2018 14:55:01",
          "content": "<p>Thank you, makes perfect sense.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 355056,
      "author_name": "edwinst",
      "author_url": "",
      "post_date": "07/10/2018 20:46:11",
      "content": "<blockquote>\n  <p>And if they can enter, then why aren't we seeing any? I also hope they don' t wait for the last minute to submit.</p>\n</blockquote>\n\n<p>One can make some hypotheses, mostly pessimistic ones: Maybe the professionals do not think there is any original approach which has a chance against the established ones. In that case a win would not be academically or scientifically valuable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 355068,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "07/10/2018 21:29:34",
          "content": "<p>A positive spin is that the professionals do not think an original approach coming from their own community is likely to emerge in a few months. Professionals also know that developing software for a new detector (like the one we simulated for the challenge, which was designed specially for the challenge, nobody has seen it before, even though it mimics broadly ATLAS and CMS future tracker ) takes years rather than months. For example we are all stunned that DBSCAN which we have never ever seen used in tracking, which we provided just because it was an almost one-liner, could be carried and improved to such significant score. We would have bet more on Hough transform which has been completely ignored, as far as we can tell from the forum. We're very impressed by the performances reached and we're looking forward to have detailed on the algorithm used.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355193,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/11/2018 06:59:43",
          "content": "<p>I don't think hough transform is ignored here.  It is implemented via 'helix unrolling', but it amounts to the same AFAIK.</p>\n\n<blockquote>\n  <p>it mimics broadly ATLAS and CMS future tracker</p>\n</blockquote>\n\n<p>Thanks for confirming that this data is different from what researchers in the field have been seeing so far.  It was said in some of the shared material, but people did not notice it really.  This gives some hope for participants I guess.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355210,
          "author_name": "edwinst",
          "author_url": "",
          "post_date": "07/11/2018 07:52:46",
          "content": "<blockquote>\n  <p>I don't think hough transform is ignored here. It is implemented via 'helix unrolling',</p>\n</blockquote>\n\n<p>I guess what David is referring to is the histogramming technique, not only the transform. My take on it would be that histogramming approaches are ignored for good reasons, but I may be wrong.</p>\n\n<p>As you imply, probably every single serious competitor is doing <em>some</em> kind of coordinate transform in the broadest sense to make the problem approachable. One might then say that massive helix unrolling combined with a merging strategy is a kind of hidden histogram approach, depending on how the merging works exactly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355214,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/11/2018 08:03:02",
          "content": "<p>I view histogramming as one way to do clustering ;)  The point is that using some polar coordinates and transforming them is used by many here, and this is the general idea of hough transform.  I agree that the specific way hough transform is used in one of the shared kernels may not be reused much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355252,
          "author_name": "trian2018",
          "author_url": "",
          "post_date": "07/11/2018 09:45:37",
          "content": "<p>It also seems to me that \"dbscan &amp; helix-unrolling\" is mainly equivalent to Hough transform. I preferred Hough transform at the beginning, because it seems intuitively right. But after quite some work, I turned to the \"helix-unrolling\"-code, which worked better, and seems to have the same substance:</p>\n\n<p>Dbscan takes the role of putting the hits into bins (as done via the \"ComboDigi\" column at the Hough transform code).</p>\n\n<p>The \"helix-unrolling\" seems equivalent to what Hough transform is really about. Each iteration at \"helix-unrolling\" does a particular transformation to the hits. Each iteration yields a particular set of long tracks: tracks which have a particular speed along the z-axis (via \"dz\"), or (e.g. if another specific main feature is used) whose mid-points, when projected onto the x-y plane, have a particular angle. I think the Hough transform code does the latter. </p>\n\n<p>Those methods, as well as the main feature (which is adjusted via the iteration variable), which is used, seem to be equivalent, but hyper-parameter optimization and computational speed might be (much) better for one method (or main feature) vs another.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355282,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "07/11/2018 11:29:29",
          "content": "<p>@Edwin, found this blog post about this competition, so I guess you're a German speaker (too). \n<a href=\"https://motls.blogspot.com/2018/07/our-edwin-steiner-is-current-leader-in.html\">https://motls.blogspot.com/2018/07/our-edwin-steiner-is-current-leader-in.html</a></p>\n\n<pre><code>Even if Outrunner improves his or her score, it's not the end of the hopes because, as Edwin says, \"Es kommt der Tag\". Or, more precisely, \"Aber es ist noch nicht aller Tage Abend!\"\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355469,
          "author_name": "edwinst",
          "author_url": "",
          "post_date": "07/11/2018 18:36:55",
          "content": "<p>@CPMP,</p>\n\n<blockquote>\n  <p>I view histogramming as one way to do clustering</p>\n</blockquote>\n\n<p>Makes sense, but the clustering only covers two dimensions of the track parameter histogram, doesn't it? The other dimensions then need to be dealt with by ensembling and merging clustering results in such approaches, for example.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355572,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/12/2018 02:29:09",
          "content": "<blockquote>\n  <p>clustering only covers two dimensions of the track parameter histogram, doesn't it?</p>\n</blockquote>\n\n<p>I am not sure I get what you mean.  You can do clustering with as many dimensions you want.  I'm using 3 dimensions for instance so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355723,
          "author_name": "edwinst",
          "author_url": "",
          "post_date": "07/12/2018 08:57:35",
          "content": "<p>The space in which you are clustering can have any dimension, that's true. I think I've seen an example kernel using 5 dimensions. However, as long as you are clustering hits and the clustering coordinates are smooth functions of (x,y,z), the points will lie on a 3-dimensional manifold in the higher-dimensional space. (What I'm saying here does not apply to clustering of more complicated derived objects, which has been suggested in at least on discussion thread.) Typically you will want to project one dimension away either completely or approximately because the separation of hits along the trajectory should not keep them from being clustered together. That's what the example kernels do: They effectively project the points onto a 2-dimensional manifold in e.g. 5-dimensional metric space. In the end, it is equivalent to a 2-dimensional clustering problem with a custom distance function. That's what I meant.</p>\n\n<p>All or some of what I said may not apply if you use clustering in a more sophisticated or indirect way. I'm referring to the approaches put forward in the public kernels.</p>\n\n<p>Note: What I call \"projection\" here may be highly non-linear functions. I do not mean linear projection only.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355792,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/12/2018 11:20:28",
          "content": "<p>Agreed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 355799,
      "author_name": "droussea",
      "author_url": "",
      "post_date": "07/12/2018 11:32:19",
      "content": "<p>Hough transform is more than helix unrolling. As indicated in <a href=\"https://www.kaggle.com/mikhailhushchyn/hough-transform\">https://www.kaggle.com/mikhailhushchyn/hough-transform</a> the key thing is that each point becomes a curve in the image space (2 dimensions in the kernel, but could have as many dimension as there are degrees of freedom in the object to be found), then the clustering is done in the image space (looking for intersection of the curves), before going back to the normal space.  While I have the impression that the evolved DBSCAN techniques mentioned in the forum are always clustering in some transformation of the normal space (where a point is always a point).  I understand this is done multiple times with variant of the transformation, then tracks found are ensembled.   </p>",
      "votes": null,
      "replies": [
        {
          "id": 355811,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/12/2018 11:56:35",
          "content": "<p>You're right, my bad.   I'm actually using a similar approach in one step of my approach, i.e. moving hits to curves in some space and finding intersections.  I missed that Hough transform was similar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355813,
          "author_name": "trian2018",
          "author_url": "",
          "post_date": "07/12/2018 11:58:27",
          "content": "<p>I think I do not agree that it is \"more\". In my understanding: At \"helix unrolling\" each hit also becomes a curve (the collection of hits you get by transforming this hit for every different angle). Then you also look for intersection of those curves (=for one specific angle, you find 10 different hits whose hit transformation for this angle lie in a small cluster).</p>\n\n<p>Anyway, this discussion is more because of intellectual curiosity I guess ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355818,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "07/12/2018 12:09:00",
          "content": "<p>OK as you@train2018  explain it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 356321,
          "author_name": "dinuuu",
          "author_url": "",
          "post_date": "07/13/2018 10:39:26",
          "content": "<p>@CPMP  Your work was too supportive sir. I have learnt a lot of things from your kernels and Your topics Sir.   Can i be an Collaborator for your team. So that i can learn some things from you sir</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 356544,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/13/2018 20:48:10",
          "content": "<p>@dineshbarri_1997, thanks a lot for your kind words.  I think I'll try to see how far I can go without teaming here.  Unlike Grzegorz, I still hope that a non specialist can fare well here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 356550,
          "author_name": "dinuuu",
          "author_url": "",
          "post_date": "07/13/2018 20:52:39",
          "content": "<p>Thank you so muchh sir. . Its privileged talking with you sir </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "354861": "Recent comments made me wonder about CERN members joining late and crushing the LB.  In most, if not all Kaggle competitions I entered, people working for the sponsor or for an affiliated entity of the sponsor cannot enter the competition or if they can, they are not eligible to prizes.  The relevant rule is 2.B but it is not clear at all to me.   Seems associated members are allowed to enter and claim prize.  I hope they don't have access to the test data and/or the code of the simulator used to generate the events.  \n\nAnd if they can enter, then why aren't we seeing any?  I also hope they don' t wait for the last minute to submit.",
    "354939": "Hi, CERN employees can participate but cannot claim any price. CERN being the largest particle physics laboratory in the world, more than half of all professional particle physicists have the so-called \"CERN user\" status, these people are not excluded. This would be equivalent for e.g. a DNA labelling competition to exclude most people with a molecular biology PhD.    However, for sure, these people (nor CERN employees in fact) have had no privileged access to the master simulation code  nor the test data. The only people to have access to the simulator are the organizers. And to be extra safe, only two of them have access to the test data truth.",
    "354956": "Thank you, makes perfect sense.",
    "355056": "&gt; And if they can enter, then why aren't we seeing any? I also hope they don' t wait for the last minute to submit.\n\nOne can make some hypotheses, mostly pessimistic ones: Maybe the professionals do not think there is any original approach which has a chance against the established ones. In that case a win would not be academically or scientifically valuable.",
    "355068": "A positive spin is that the professionals do not think an original approach coming from their own community is likely to emerge in a few months. Professionals also know that developing software for a new detector (like the one we simulated for the challenge, which was designed specially for the challenge, nobody has seen it before, even though it mimics broadly ATLAS and CMS future tracker ) takes years rather than months. For example we are all stunned that DBSCAN which we have never ever seen used in tracking, which we provided just because it was an almost one-liner, could be carried and improved to such significant score. We would have bet more on Hough transform which has been completely ignored, as far as we can tell from the forum. We're very impressed by the performances reached and we're looking forward to have detailed on the algorithm used.",
    "355193": "I don't think hough transform is ignored here.  It is implemented via 'helix unrolling', but it amounts to the same AFAIK.\n\n&gt; it mimics broadly ATLAS and CMS future tracker\n\nThanks for confirming that this data is different from what researchers in the field have been seeing so far.  It was said in some of the shared material, but people did not notice it really.  This gives some hope for participants I guess.",
    "355210": "&gt; I don't think hough transform is ignored here. It is implemented via 'helix unrolling',\n\nI guess what David is referring to is the histogramming technique, not only the transform. My take on it would be that histogramming approaches are ignored for good reasons, but I may be wrong.\n\nAs you imply, probably every single serious competitor is doing *some* kind of coordinate transform in the broadest sense to make the problem approachable. One might then say that massive helix unrolling combined with a merging strategy is a kind of hidden histogram approach, depending on how the merging works exactly.",
    "355214": "I view histogramming as one way to do clustering ;)  The point is that using some polar coordinates and transforming them is used by many here, and this is the general idea of hough transform.  I agree that the specific way hough transform is used in one of the shared kernels may not be reused much.",
    "355252": "It also seems to me that \"dbscan &amp; helix-unrolling\" is mainly equivalent to Hough transform. I preferred Hough transform at the beginning, because it seems intuitively right. But after quite some work, I turned to the \"helix-unrolling\"-code, which worked better, and seems to have the same substance:\n\nDbscan takes the role of putting the hits into bins (as done via the \"ComboDigi\" column at the Hough transform code).\n\nThe \"helix-unrolling\" seems equivalent to what Hough transform is really about. Each iteration at \"helix-unrolling\" does a particular transformation to the hits. Each iteration yields a particular set of long tracks: tracks which have a particular speed along the z-axis (via \"dz\"), or (e.g. if another specific main feature is used) whose mid-points, when projected onto the x-y plane, have a particular angle. I think the Hough transform code does the latter. \n\nThose methods, as well as the main feature (which is adjusted via the iteration variable), which is used, seem to be equivalent, but hyper-parameter optimization and computational speed might be (much) better for one method (or main feature) vs another.",
    "355282": "Edwin, found this blog post about this competition, so I guess you're a German speaker (too). \nhttps://motls.blogspot.com/2018/07/our-edwin-steiner-is-current-leader-in.html\n\n    Even if Outrunner improves his or her score, it's not the end of the hopes because, as Edwin says, \"Es kommt der Tag\". Or, more precisely, \"Aber es ist noch nicht aller Tage Abend!\"",
    "355469": "CPMP,\n\n&gt; I view histogramming as one way to do clustering\n\nMakes sense, but the clustering only covers two dimensions of the track parameter histogram, doesn't it? The other dimensions then need to be dealt with by ensembling and merging clustering results in such approaches, for example.",
    "355572": "&gt; clustering only covers two dimensions of the track parameter histogram, doesn't it?\n\nI am not sure I get what you mean.  You can do clustering with as many dimensions you want.  I'm using 3 dimensions for instance so far.",
    "355723": "The space in which you are clustering can have any dimension, that's true. I think I've seen an example kernel using 5 dimensions. However, as long as you are clustering hits and the clustering coordinates are smooth functions of (x,y,z), the points will lie on a 3-dimensional manifold in the higher-dimensional space. (What I'm saying here does not apply to clustering of more complicated derived objects, which has been suggested in at least on discussion thread.) Typically you will want to project one dimension away either completely or approximately because the separation of hits along the trajectory should not keep them from being clustered together. That's what the example kernels do: They effectively project the points onto a 2-dimensional manifold in e.g. 5-dimensional metric space. In the end, it is equivalent to a 2-dimensional clustering problem with a custom distance function. That's what I meant.\n\nAll or some of what I said may not apply if you use clustering in a more sophisticated or indirect way. I'm referring to the approaches put forward in the public kernels.\n\nNote: What I call \"projection\" here may be highly non-linear functions. I do not mean linear projection only.",
    "355792": "Agreed.",
    "355799": "Hough transform is more than helix unrolling. As indicated in https://www.kaggle.com/mikhailhushchyn/hough-transform the key thing is that each point becomes a curve in the image space (2 dimensions in the kernel, but could have as many dimension as there are degrees of freedom in the object to be found), then the clustering is done in the image space (looking for intersection of the curves), before going back to the normal space.  While I have the impression that the evolved DBSCAN techniques mentioned in the forum are always clustering in some transformation of the normal space (where a point is always a point).  I understand this is done multiple times with variant of the transformation, then tracks found are ensembled.",
    "355811": "You're right, my bad.   I'm actually using a similar approach in one step of my approach, i.e. moving hits to curves in some space and finding intersections.  I missed that Hough transform was similar.",
    "355813": "I think I do not agree that it is \"more\". In my understanding: At \"helix unrolling\" each hit also becomes a curve (the collection of hits you get by transforming this hit for every different angle). Then you also look for intersection of those curves (=for one specific angle, you find 10 different hits whose hit transformation for this angle lie in a small cluster).\n\nAnyway, this discussion is more because of intellectual curiosity I guess ;)",
    "355818": "OK as you@train2018  explain it.",
    "356321": "CPMP  Your work was too supportive sir. I have learnt a lot of things from your kernels and Your topics Sir.   Can i be an Collaborator for your team. So that i can learn some things from you sir",
    "356544": "dineshbarri_1997, thanks a lot for your kind words.  I think I'll try to see how far I can go without teaming here.  Unlike Grzegorz, I still hope that a non specialist can fare well here.",
    "356550": "Thank you so muchh sir. . Its privileged talking with you sir"
  },
  "source": "meta"
}