{
  "id": 61590,
  "title": "Criteria for good features ",
  "url": "/competitions/trackml-particle-identification/discussion/61590",
  "author_name": "yuval reina",
  "post_date": "2018-07-21T09:09:40.169000",
  "votes": 25,
  "comment_count": 27,
  "views": 0,
  "content": "<p>I wanted to share my views about the criteria to select good features for this competition (most of what I’m going to write already appear as comments in other discussions)</p>\n\n<p>My criteria are:</p>\n\n<ol>\n<li>The feature must be constant in an ideal track.</li>\n<li>The feature (or ensemble of features) should get different values for different tracks</li>\n<li>The feature must be bounded</li>\n<li>The feature should be distributed as uniformly as possible</li>\n<li>less is better - This is not a feature engineering competition, we need to find the minimal number of geometrical/physical oriented criteria that give a good description of the tracks  </li>\n</ol>\n\n<p>As an example for a problematic feature I’ll take a feature that is commonly used but fail to meet three of these criteria: <strong>z/r</strong>.</p>\n\n<p>This feature is not bounded and is not uniformly distributed. Tracks that have high velocity in Z would have very large z/r values, and tracks with low Z velocity would get very low value. When we get a 0.1 difference between two z/r values, how do you decide if it is big or small?</p>\n\n<p>An even bigger problem with this feature is that it is not constant for a single helix because it does not consider the arc the track is following, it just assumes the particle goes straight from (0,0) to (1,1).</p>\n\n<p>One can solve the first two issues by using arctan(z/r), for the 3rd issue a little bit of geometry is needed  </p>",
  "messages": [
    {
      "id": 360007,
      "postDate": "2018-07-21T09:09:40.170Z",
      "content": "<p>I wanted to share my views about the criteria to select good features for this competition (most of what I’m going to write already appear as comments in other discussions)</p>\n\n<p>My criteria are:</p>\n\n<ol>\n<li>The feature must be constant in an ideal track.</li>\n<li>The feature (or ensemble of features) should get different values for different tracks</li>\n<li>The feature must be bounded</li>\n<li>The feature should be distributed as uniformly as possible</li>\n<li>less is better - This is not a feature engineering competition, we need to find the minimal number of geometrical/physical oriented criteria that give a good description of the tracks  </li>\n</ol>\n\n<p>As an example for a problematic feature I’ll take a feature that is commonly used but fail to meet three of these criteria: <strong>z/r</strong>.</p>\n\n<p>This feature is not bounded and is not uniformly distributed. Tracks that have high velocity in Z would have very large z/r values, and tracks with low Z velocity would get very low value. When we get a 0.1 difference between two z/r values, how do you decide if it is big or small?</p>\n\n<p>An even bigger problem with this feature is that it is not constant for a single helix because it does not consider the arc the track is following, it just assumes the particle goes straight from (0,0) to (1,1).</p>\n\n<p>One can solve the first two issues by using arctan(z/r), for the 3rd issue a little bit of geometry is needed  </p>",
      "rawMarkdown": "I wanted to share my views about the criteria to select good features for this competition (most of what I’m going to write already appear as comments in other discussions)\n\n My criteria are:\n\n 1.  The feature must be constant in an ideal track.\n 2. The feature (or ensemble of features) should get different values for different tracks\n 3. The feature must be bounded\n 4. The feature should be distributed as uniformly as possible\n 5. less is better - This is not a feature engineering competition, we need to find the minimal number of geometrical/physical oriented criteria that give a good description of the tracks  \n\nAs an example for a problematic feature I’ll take a feature that is commonly used but fail to meet three of these criteria: **z/r**.\n\nThis feature is not bounded and is not uniformly distributed. Tracks that have high velocity in Z would have very large z/r values, and tracks with low Z velocity would get very low value. When we get a 0.1 difference between two z/r values, how do you decide if it is big or small?\n\nAn even bigger problem with this feature is that it is not constant for a single helix because it does not consider the arc the track is following, it just assumes the particle goes straight from (0,0) to (1,1).\n\nOne can solve the first two issues by using arctan(z/r), for the 3rd issue a little bit of geometry is needed  \n",
      "votes": 25
    },
    {
      "id": 360169,
      "postDate": "2018-07-21T19:07:55.213Z",
      "content": "<p>z/r is bounded unless mistaken, because r is lower bounded above 0, somewhere around 30 mm.  But it is certainly not uniformly distributed.</p>",
      "rawMarkdown": "z/r is bounded unless mistaken, because r is lower bounded above 0, somewhere around 30 mm.  But it is certainly not uniformly distributed.",
      "votes": 1
    },
    {
      "id": 360376,
      "postDate": "2018-07-22T10:04:34.073Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 360385,
          "postDate": "2018-07-22T11:10:47.147Z",
          "content": "<p><a href=\"/starhao\">@starhao</a> Only commenting out some of the features wouldn't do the trick. You need to construct new features that will represent the Helix properly and efficiently. </p>",
          "rawMarkdown": "@starhao Only commenting out some of the features wouldn't do the trick. You need to construct new features that will represent the Helix properly and efficiently. ",
          "votes": 4
        },
        {
          "id": 360452,
          "postDate": "2018-07-22T13:43:42.820Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 360473,
          "postDate": "2018-07-22T14:38:51.160Z",
          "content": "<p>Yes, the arctan fixes the unbounded (i.e has a very large dynamic range) and uniformity problem.\nAs for the next comment, imagine the Helix as a spiral staircase, where is the constant slope? </p>",
          "rawMarkdown": "Yes, the arctan fixes the unbounded (i.e has a very large dynamic range) and uniformity problem.\nAs for the next comment, imagine the Helix as a spiral staircase, where is the constant slope? ",
          "votes": 2
        },
        {
          "id": 360476,
          "postDate": "2018-07-22T14:49:39.650Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 365211,
          "postDate": "2018-08-02T05:50:12.413Z",
          "content": "<p>@yuval. The constant slope is the angle formed by the tangent line at each point with the z axis. If we use only z, r, (and all information we have: x, y), I'm afraid we are not able to solve the constant problem simply like solving the uniformity using arctan(z/r)</p>",
          "rawMarkdown": "@yuval. The constant slope is the angle formed by the tangent line at each point with the z axis. If we use only z, r, (and all information we have: x, y), I'm afraid we are not able to solve the constant problem simply like solving the uniformity using arctan(z/r)",
          "votes": 1
        },
        {
          "id": 365221,
          "postDate": "2018-08-02T06:17:32.843Z",
          "content": "<p>You are getting warmer. What about the integral?</p>\n\n<p>The one MOST impotent criteria for a feature is - it must be <strong>constant</strong> for a <strong>perfect known helix</strong> which starts at the origin.  perfect = mathematical helix. constant = to the accuracy of floating point calculations.  </p>",
          "rawMarkdown": "You are getting warmer. What about the integral?\n\nThe one MOST impotent criteria for a feature is - it must be **constant** for a **perfect known helix** which starts at the origin.  perfect = mathematical helix. constant = to the accuracy of floating point calculations.  ",
          "votes": 4
        },
        {
          "id": 366549,
          "postDate": "2018-08-05T19:10:33.383Z",
          "content": "<p>I don't think I've got good features for all directions / range of momentums - I've got some that seem to work well for certain behaviours of particles but seemingly pretty badly for others. \nDo you just use the same 2 features all the time in your binning approach?</p>",
          "rawMarkdown": "I don't think I've got good features for all directions / range of momentums - I've got some that seem to work well for certain behaviours of particles but seemingly pretty badly for others. \nDo you just use the same 2 features all the time in your binning approach?"
        },
        {
          "id": 366550,
          "postDate": "2018-08-05T19:25:32.283Z",
          "content": "<p>I don't know for yuval,  but I'm using the same 2 features for dbscan (2 if you count cos and sin as one).  Loop of dbscan with merging on the fly yields a local score that is getting close to 0.76 with these 2 features.</p>",
          "rawMarkdown": "I don't know for yuval,  but I'm using the same 2 features for dbscan (2 if you count cos and sin as one).  Loop of dbscan with merging on the fly yields a local score that is getting close to 0.76 with these 2 features.",
          "votes": 1
        },
        {
          "id": 366552,
          "postDate": "2018-08-05T19:36:48.830Z",
          "content": "<p>Thanks, presumably that means you are capturing most (90+%) of particles that cross the z axis near the origin (I think this limit was .82 as demonstrated in a kernel)\nLooks like I need a better set of features then!</p>",
          "rawMarkdown": "Thanks, presumably that means you are capturing most (90+%) of particles that cross the z axis near the origin (I think this limit was .82 as demonstrated in a kernel)\nLooks like I need a better set of features then!"
        },
        {
          "id": 367866,
          "postDate": "2018-08-08T17:46:36.663Z",
          "content": "<p>I got stuck at the second feature. I arrived at the same conclusion as Kha A. Vo, I'm afraid there is not enough information to figure out what the slope between XY and Z is, or am I missing something?</p>",
          "rawMarkdown": "I got stuck at the second feature. I arrived at the same conclusion as Kha A. Vo, I'm afraid there is not enough information to figure out what the slope between XY and Z is, or am I missing something?"
        },
        {
          "id": 367881,
          "postDate": "2018-08-08T18:30:00.267Z",
          "content": "<p>You are missing - the slop is the distance traveled by the particle in the XY plane divided by Z. for a specific XY  you need one assumption (which you also need for the other feature) to calculate the distance.  </p>",
          "rawMarkdown": "You are missing - the slop is the distance traveled by the particle in the XY plane divided by Z. for a specific XY  you need one assumption (which you also need for the other feature) to calculate the distance.  ",
          "votes": 2
        },
        {
          "id": 367916,
          "postDate": "2018-08-08T20:19:27.883Z",
          "content": "<p>Thanks for the fast reply. This much I understood, to be more precise dalpha/dz should be the constant (tangencial slope), what I didn't get is how to get dalpha and dz from the hits list. I thought about setting dz as another variable to cycle through. But then I would end up with 1/r0, z0 and dz to loop through, which I doubt was the case. What I ended up using that give's me a score of .17 is arctan2((z-z0),r). I think I'm overlooking something very simple in my assumption.</p>",
          "rawMarkdown": "Thanks for the fast reply. This much I understood, to be more precise dalpha/dz should be the constant (tangencial slope), what I didn't get is how to get dalpha and dz from the hits list. I thought about setting dz as another variable to cycle through. But then I would end up with 1/r0, z0 and dz to loop through, which I doubt was the case. What I ended up using that give's me a score of .17 is arctan2((z-z0),r). I think I'm overlooking something very simple in my assumption."
        },
        {
          "id": 367920,
          "postDate": "2018-08-08T20:28:29.687Z",
          "content": "<p>When you take Z-Z0, this is the distance the particle traveled in the Z direction. Is <code>r</code>  the distance the particle traveled in the XY plan? </p>\n\n<p>in dalpha/dZ you have a units problem - the units are rad/mm and the slope has no units.  </p>",
          "rawMarkdown": "When you take Z-Z0, this is the distance the particle traveled in the Z direction. Is `r`  the distance the particle traveled in the XY plan? \n\nin dalpha/dZ you have a units problem - the units are rad/mm and the slope has no units.  "
        },
        {
          "id": 368070,
          "postDate": "2018-08-09T07:48:57.030Z",
          "content": "<p>@yuval Is your number of bins for each feature roughly the same in your method? </p>",
          "rawMarkdown": "@yuval Is your number of bins for each feature roughly the same in your method? "
        },
        {
          "id": 368074,
          "postDate": "2018-08-09T07:54:12.947Z",
          "content": "<p>One has twice the number bins. Although because I do sparse binning, I don't really count the number of bins for each features</p>",
          "rawMarkdown": "One has twice the number bins. Although because I do sparse binning, I don't really count the number of bins for each features"
        },
        {
          "id": 368077,
          "postDate": "2018-08-09T08:04:24.393Z",
          "content": "<p>@yuval. I really would like to know your performance after around 1000 first loops. My binning now gets to roughly 0.25 after 1000 loops, and after that it reaches the plateau around 0.3. You said before that you reach 0.63 with 5500 loops?! Then for the first 1000 did you obtain 0.4~0.5? Did your score increasing uniformly from the beginning and after a certain stage, it decelerates/ or it decelerates from the beginning?</p>",
          "rawMarkdown": "@yuval. I really would like to know your performance after around 1000 first loops. My binning now gets to roughly 0.25 after 1000 loops, and after that it reaches the plateau around 0.3. You said before that you reach 0.63 with 5500 loops?! Then for the first 1000 did you obtain 0.4~0.5? Did your score increasing uniformly from the beginning and after a certain stage, it decelerates/ or it decelerates from the beginning?"
        },
        {
          "id": 368083,
          "postDate": "2018-08-09T08:22:40.970Z",
          "content": "<p>I'm also very interested in this. I'm stuck at .17, still tuning the parameters. Btw, I'm looping f1 50 times and f2 20.</p>",
          "rawMarkdown": "I'm also very interested in this. I'm stuck at .17, still tuning the parameters. Btw, I'm looping f1 50 times and f2 20."
        },
        {
          "id": 368105,
          "postDate": "2018-08-09T09:07:48.883Z",
          "content": "<p>After 500 I get 0.31. \nThe score does not increase uniformly. It increase with decreasing slope.</p>",
          "rawMarkdown": "After 500 I get 0.31. \nThe score does not increase uniformly. It increase with decreasing slope.",
          "votes": 3
        },
        {
          "id": 368330,
          "postDate": "2018-08-09T17:54:05.437Z",
          "content": "<p>I found out I'm still struggling with the second feature. It's constant only in a specific interval, it breaks symmetry outside of that.</p>",
          "rawMarkdown": "I found out I'm still struggling with the second feature. It's constant only in a specific interval, it breaks symmetry outside of that."
        },
        {
          "id": 368389,
          "postDate": "2018-08-09T19:55:52.297Z",
          "content": "<p>I don't know what you mean by specific interval. Be aware that if the arc in XY is longer then pi/2 (angle wise) you need a special treatment for the features (because you need to solve the ambiguity). fortunately there are only few tracks like this, and this issue can be ignored (until your scoe is high enough)\nIf this is not  what you mean by \"specific intervals\"  then you'll have to look at your feature again</p>",
          "rawMarkdown": "I don't know what you mean by specific interval. Be aware that if the arc in XY is longer then pi/2 (angle wise) you need a special treatment for the features (because you need to solve the ambiguity). fortunately there are only few tracks like this, and this issue can be ignored (until your scoe is high enough)\nIf this is not  what you mean by \"specific intervals\"  then you'll have to look at your feature again",
          "votes": 1
        },
        {
          "id": 368400,
          "postDate": "2018-08-09T20:29:40.560Z",
          "content": "<p>This subthread helps me a lot. I almost think I've got the crux of your features, but I haven't T_T. </p>\n\n<p>My issue is the second feature. It is constant on a perfect helix (as the one in your other thread about testing whether features are good or not), but it is not constant on some real tracks and it actually can vary a lot. I think the problem is that some tracks have very small velocity in the z direction, so the accuracy of the second feature suffers from the energy dissipated by the detector.</p>\n\n<p>How do you handle this?</p>",
          "rawMarkdown": "This subthread helps me a lot. I almost think I've got the crux of your features, but I haven't T_T. \n\nMy issue is the second feature. It is constant on a perfect helix (as the one in your other thread about testing whether features are good or not), but it is not constant on some real tracks and it actually can vary a lot. I think the problem is that some tracks have very small velocity in the z direction, so the accuracy of the second feature suffers from the energy dissipated by the detector.\n\nHow do you handle this?"
        },
        {
          "id": 368416,
          "postDate": "2018-08-09T21:16:30.230Z",
          "content": "<p>I think I'm getting there. Just managed to get to .31 with 500 iterations, but with 5000 I'm stuck at .38.</p>",
          "rawMarkdown": "I think I'm getting there. Just managed to get to .31 with 500 iterations, but with 5000 I'm stuck at .38."
        },
        {
          "id": 368448,
          "postDate": "2018-08-09T22:39:40.020Z",
          "content": "<p>You may have clustered too much. If your clustering algorithm+meta parameters are too wide you may cluster too many hits together. When you do this, the score may rise fast at the beginning and then plateau too early.\nYou can measure it if you calculate the score for different tracks and then sum the weighs of hit in this track.\nWhen I do this calculation for long tracks (10 hits and above) I get more then 98% accuracy (score/total weight) how much do you get?</p>",
          "rawMarkdown": "You may have clustered too much. If your clustering algorithm+meta parameters are too wide you may cluster too many hits together. When you do this, the score may rise fast at the beginning and then plateau too early.\nYou can measure it if you calculate the score for different tracks and then sum the weighs of hit in this track.\nWhen I do this calculation for long tracks (10 hits and above) I get more then 98% accuracy (score/total weight) how much do you get?",
          "votes": 3
        },
        {
          "id": 368763,
          "postDate": "2018-08-10T17:59:32.423Z",
          "content": "<p>I didn't try that, but I don't think I'm clustering many hits together. I plotted the data in a scatter plot and estimated the minimum number of bins for my 2 features. I have about 4x that number now, currently the best configuration I found so far. With that I get .32 on 450 loops on 1 feature only. More loops have diminishing returns. Looping through my second feature doesn't produce much better results, I think that's the problem... I need to find the sweet spot for the second feature.</p>\n\n<p>EDIT: I just tried what you suggested, I get 62% of the tracks I scored with 10 or more hits, that's not very high... Need to play more with the number of bins then. Thanks for all the help so far</p>",
          "rawMarkdown": "I didn't try that, but I don't think I'm clustering many hits together. I plotted the data in a scatter plot and estimated the minimum number of bins for my 2 features. I have about 4x that number now, currently the best configuration I found so far. With that I get .32 on 450 loops on 1 feature only. More loops have diminishing returns. Looping through my second feature doesn't produce much better results, I think that's the problem... I need to find the sweet spot for the second feature.\n\nEDIT: I just tried what you suggested, I get 62% of the tracks I scored with 10 or more hits, that's not very high... Need to play more with the number of bins then. Thanks for all the help so far"
        }
      ]
    },
    {
      "id": 360069,
      "postDate": "2018-07-21T13:00:37.527Z",
      "content": "<p>Thanks @yuval r. </p>",
      "rawMarkdown": "Thanks @yuval r. "
    }
  ],
  "comments": [
    {
      "id": 360169,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-07-21T19:07:55.213000",
      "content": "<p>z/r is bounded unless mistaken, because r is lower bounded above 0, somewhere around 30 mm.  But it is certainly not uniformly distributed.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 360376,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-22T10:04:34.073000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 360385,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-22T11:10:47.147000",
          "content": "<p><a href=\"/starhao\">@starhao</a> Only commenting out some of the features wouldn't do the trick. You need to construct new features that will represent the Helix properly and efficiently. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 360452,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-22T13:43:42.820000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 360473,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-07-22T14:38:51.160000",
          "content": "<p>Yes, the arctan fixes the unbounded (i.e has a very large dynamic range) and uniformity problem.\nAs for the next comment, imagine the Helix as a spiral staircase, where is the constant slope? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 360476,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-07-22T14:49:39.650000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 365211,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-08-02T05:50:12.413000",
          "content": "<p>@yuval. The constant slope is the angle formed by the tangent line at each point with the z axis. If we use only z, r, (and all information we have: x, y), I'm afraid we are not able to solve the constant problem simply like solving the uniformity using arctan(z/r)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 365221,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-02T06:17:32.843000",
          "content": "<p>You are getting warmer. What about the integral?</p>\n\n<p>The one MOST impotent criteria for a feature is - it must be <strong>constant</strong> for a <strong>perfect known helix</strong> which starts at the origin.  perfect = mathematical helix. constant = to the accuracy of floating point calculations.  </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 366549,
          "author_name": "Seb",
          "author_url": "",
          "post_date": "2018-08-05T19:10:33.383000",
          "content": "<p>I don't think I've got good features for all directions / range of momentums - I've got some that seem to work well for certain behaviours of particles but seemingly pretty badly for others. \nDo you just use the same 2 features all the time in your binning approach?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366550,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-05T19:25:32.283000",
          "content": "<p>I don't know for yuval,  but I'm using the same 2 features for dbscan (2 if you count cos and sin as one).  Loop of dbscan with merging on the fly yields a local score that is getting close to 0.76 with these 2 features.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 366552,
          "author_name": "Seb",
          "author_url": "",
          "post_date": "2018-08-05T19:36:48.830000",
          "content": "<p>Thanks, presumably that means you are capturing most (90+%) of particles that cross the z axis near the origin (I think this limit was .82 as demonstrated in a kernel)\nLooks like I need a better set of features then!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367866,
          "author_name": "Eduardo Konishi",
          "author_url": "",
          "post_date": "2018-08-08T17:46:36.663000",
          "content": "<p>I got stuck at the second feature. I arrived at the same conclusion as Kha A. Vo, I'm afraid there is not enough information to figure out what the slope between XY and Z is, or am I missing something?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367881,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-08T18:30:00.267000",
          "content": "<p>You are missing - the slop is the distance traveled by the particle in the XY plane divided by Z. for a specific XY  you need one assumption (which you also need for the other feature) to calculate the distance.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 367916,
          "author_name": "Eduardo Konishi",
          "author_url": "",
          "post_date": "2018-08-08T20:19:27.883000",
          "content": "<p>Thanks for the fast reply. This much I understood, to be more precise dalpha/dz should be the constant (tangencial slope), what I didn't get is how to get dalpha and dz from the hits list. I thought about setting dz as another variable to cycle through. But then I would end up with 1/r0, z0 and dz to loop through, which I doubt was the case. What I ended up using that give's me a score of .17 is arctan2((z-z0),r). I think I'm overlooking something very simple in my assumption.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367920,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-08T20:28:29.687000",
          "content": "<p>When you take Z-Z0, this is the distance the particle traveled in the Z direction. Is <code>r</code>  the distance the particle traveled in the XY plan? </p>\n\n<p>in dalpha/dZ you have a units problem - the units are rad/mm and the slope has no units.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 368070,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-08-09T07:48:57.030000",
          "content": "<p>@yuval Is your number of bins for each feature roughly the same in your method? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 368074,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-09T07:54:12.947000",
          "content": "<p>One has twice the number bins. Although because I do sparse binning, I don't really count the number of bins for each features</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 368077,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2018-08-09T08:04:24.393000",
          "content": "<p>@yuval. I really would like to know your performance after around 1000 first loops. My binning now gets to roughly 0.25 after 1000 loops, and after that it reaches the plateau around 0.3. You said before that you reach 0.63 with 5500 loops?! Then for the first 1000 did you obtain 0.4~0.5? Did your score increasing uniformly from the beginning and after a certain stage, it decelerates/ or it decelerates from the beginning?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 368083,
          "author_name": "Eduardo Konishi",
          "author_url": "",
          "post_date": "2018-08-09T08:22:40.970000",
          "content": "<p>I'm also very interested in this. I'm stuck at .17, still tuning the parameters. Btw, I'm looping f1 50 times and f2 20.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 368105,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-09T09:07:48.883000",
          "content": "<p>After 500 I get 0.31. \nThe score does not increase uniformly. It increase with decreasing slope.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 368330,
          "author_name": "Eduardo Konishi",
          "author_url": "",
          "post_date": "2018-08-09T17:54:05.437000",
          "content": "<p>I found out I'm still struggling with the second feature. It's constant only in a specific interval, it breaks symmetry outside of that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 368389,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-09T19:55:52.297000",
          "content": "<p>I don't know what you mean by specific interval. Be aware that if the arc in XY is longer then pi/2 (angle wise) you need a special treatment for the features (because you need to solve the ambiguity). fortunately there are only few tracks like this, and this issue can be ignored (until your scoe is high enough)\nIf this is not  what you mean by \"specific intervals\"  then you'll have to look at your feature again</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 368400,
          "author_name": "Yang Wang",
          "author_url": "",
          "post_date": "2018-08-09T20:29:40.560000",
          "content": "<p>This subthread helps me a lot. I almost think I've got the crux of your features, but I haven't T_T. </p>\n\n<p>My issue is the second feature. It is constant on a perfect helix (as the one in your other thread about testing whether features are good or not), but it is not constant on some real tracks and it actually can vary a lot. I think the problem is that some tracks have very small velocity in the z direction, so the accuracy of the second feature suffers from the energy dissipated by the detector.</p>\n\n<p>How do you handle this?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 368416,
          "author_name": "Eduardo Konishi",
          "author_url": "",
          "post_date": "2018-08-09T21:16:30.230000",
          "content": "<p>I think I'm getting there. Just managed to get to .31 with 500 iterations, but with 5000 I'm stuck at .38.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 368448,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-09T22:39:40.020000",
          "content": "<p>You may have clustered too much. If your clustering algorithm+meta parameters are too wide you may cluster too many hits together. When you do this, the score may rise fast at the beginning and then plateau too early.\nYou can measure it if you calculate the score for different tracks and then sum the weighs of hit in this track.\nWhen I do this calculation for long tracks (10 hits and above) I get more then 98% accuracy (score/total weight) how much do you get?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 368763,
          "author_name": "Eduardo Konishi",
          "author_url": "",
          "post_date": "2018-08-10T17:59:32.423000",
          "content": "<p>I didn't try that, but I don't think I'm clustering many hits together. I plotted the data in a scatter plot and estimated the minimum number of bins for my 2 features. I have about 4x that number now, currently the best configuration I found so far. With that I get .32 on 450 loops on 1 feature only. More loops have diminishing returns. Looping through my second feature doesn't produce much better results, I think that's the problem... I need to find the sweet spot for the second feature.</p>\n\n<p>EDIT: I just tried what you suggested, I get 62% of the tracks I scored with 10 or more hits, that's not very high... Need to play more with the number of bins then. Thanks for all the help so far</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 360069,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2018-07-21T13:00:37.527000",
      "content": "<p>Thanks @yuval r. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "360007": "I wanted to share my views about the criteria to select good features for this competition (most of what I’m going to write already appear as comments in other discussions)\n\n My criteria are:\n\n 1.  The feature must be constant in an ideal track.\n 2. The feature (or ensemble of features) should get different values for different tracks\n 3. The feature must be bounded\n 4. The feature should be distributed as uniformly as possible\n 5. less is better - This is not a feature engineering competition, we need to find the minimal number of geometrical/physical oriented criteria that give a good description of the tracks  \n\nAs an example for a problematic feature I’ll take a feature that is commonly used but fail to meet three of these criteria: **z/r**.\n\nThis feature is not bounded and is not uniformly distributed. Tracks that have high velocity in Z would have very large z/r values, and tracks with low Z velocity would get very low value. When we get a 0.1 difference between two z/r values, how do you decide if it is big or small?\n\nAn even bigger problem with this feature is that it is not constant for a single helix because it does not consider the arc the track is following, it just assumes the particle goes straight from (0,0) to (1,1).\n\nOne can solve the first two issues by using arctan(z/r), for the 3rd issue a little bit of geometry is needed  \n",
    "360169": "z/r is bounded unless mistaken, because r is lower bounded above 0, somewhere around 30 mm.  But it is certainly not uniformly distributed.",
    "360376": "",
    "360069": "Thanks @yuval r. "
  }
}