{
  "id": 62804,
  "title": "Struggling With Clustering? Do this test. ",
  "url": "/competitions/trackml-particle-identification/discussion/62804",
  "author_name": "yuval reina",
  "post_date": "2018-08-07T11:03:21.916000",
  "votes": 25,
  "comment_count": 32,
  "views": 0,
  "content": "<p>This post is written  for kagglers who are using clustering as their main track detection mechanism and are still below LB of 0.7.</p>\n\n<p>If this is your situation,  probably you have an issue with your features.\nTo understand and fix this issue, do the following:</p>\n\n<ol>\n<li>Run this code:</li>\n</ol>\n\n<p>R=2000    #Helix Radius</p>\n\n<p>w=np.pi/3 #length of arc</p>\n\n<p>v=3000    # velocity in Z</p>\n\n<p>theta0=-0.3  # angle at the origin</p>\n\n<p>D=0    #D=500   -  Use for helixes far from the origin</p>\n\n<p>z0=0    #z0=200   -  Z shifting</p>\n\n<p>t=np.random.rand(50)</p>\n\n<p>X=-R*np.sin(w*t+theta0)+(R-D)*np.sin(theta0)</p>\n\n<p>Y=-R*np.cos(w*t+theta0)+(R-D)*np.cos(theta0)</p>\n\n<p>Z=z0+v*t</p>\n\n<p>Now you have a simulation of a perfect helix, starting at the origin (X,Y,Z) </p>\n\n<ol>\n<li><p>calculate your features for this helix</p></li>\n<li><p>if your features are (F1, F2, ...Fn)  and your scanning parameters are (S1,S2,..Sm), you should be able to find a set of {S} for which {F} are constant for all values of <code>t</code> </p></li>\n</ol>\n\n<p>Be ware \"<strong>constant</strong>\" is <strong>constant</strong>  not approximately the same. \nIf you can't get the features to be constant you don't have the right features\n(use as few features and scanning parameters as you can - less is more) </p>",
  "messages": [
    {
      "id": 367235,
      "postDate": "2018-08-07T11:03:21.917Z",
      "content": "<p>This post is written  for kagglers who are using clustering as their main track detection mechanism and are still below LB of 0.7.</p>\n\n<p>If this is your situation,  probably you have an issue with your features.\nTo understand and fix this issue, do the following:</p>\n\n<ol>\n<li>Run this code:</li>\n</ol>\n\n<p>R=2000    #Helix Radius</p>\n\n<p>w=np.pi/3 #length of arc</p>\n\n<p>v=3000    # velocity in Z</p>\n\n<p>theta0=-0.3  # angle at the origin</p>\n\n<p>D=0    #D=500   -  Use for helixes far from the origin</p>\n\n<p>z0=0    #z0=200   -  Z shifting</p>\n\n<p>t=np.random.rand(50)</p>\n\n<p>X=-R*np.sin(w*t+theta0)+(R-D)*np.sin(theta0)</p>\n\n<p>Y=-R*np.cos(w*t+theta0)+(R-D)*np.cos(theta0)</p>\n\n<p>Z=z0+v*t</p>\n\n<p>Now you have a simulation of a perfect helix, starting at the origin (X,Y,Z) </p>\n\n<ol>\n<li><p>calculate your features for this helix</p></li>\n<li><p>if your features are (F1, F2, ...Fn)  and your scanning parameters are (S1,S2,..Sm), you should be able to find a set of {S} for which {F} are constant for all values of <code>t</code> </p></li>\n</ol>\n\n<p>Be ware \"<strong>constant</strong>\" is <strong>constant</strong>  not approximately the same. \nIf you can't get the features to be constant you don't have the right features\n(use as few features and scanning parameters as you can - less is more) </p>",
      "rawMarkdown": "This post is written  for kagglers who are using clustering as their main track detection mechanism and are still below LB of 0.7.\n\nIf this is your situation,  probably you have an issue with your features.\nTo understand and fix this issue, do the following:\n\n1. Run this code:\n\nR=2000    #Helix Radius\n\nw=np.pi/3 #length of arc\n\nv=3000    # velocity in Z\n\ntheta0=-0.3  # angle at the origin\n\nD=0    #D=500   -  Use for helixes far from the origin\n\nz0=0    #z0=200   -  Z shifting\n\nt=np.random.rand(50)\n\nX=-R*np.sin(w*t+theta0)+(R-D)*np.sin(theta0)\n\nY=-R*np.cos(w*t+theta0)+(R-D)*np.cos(theta0)\n\nZ=z0+v*t\n\n\nNow you have a simulation of a perfect helix, starting at the origin (X,Y,Z) \n\n2. calculate your features for this helix\n\n3. if your features are (F1, F2, ...Fn)  and your scanning parameters are (S1,S2,..Sm), you should be able to find a set of {S} for which {F} are constant for all values of `t` \n\nBe ware \"**constant**\" is **constant**  not approximately the same. \nIf you can't get the features to be constant you don't have the right features\n(use as few features and scanning parameters as you can - less is more) \n",
      "votes": 25
    },
    {
      "id": 367730,
      "postDate": "2018-08-08T12:14:45.963Z",
      "content": "<p>@yuval, I admire your wilI to get others pass you ;)  Let me add two grain of salt to what you shared.</p>\n\n<ol>\n<li><p>We can relax a bit the constraint to have exactly constant parameters when we use DBSCAN, as it can cluster points that are not identical.  It means we don't need to have scanning parameters that yield absolutely constant values for a given helix.  Key is to get parameters closer to each other for a given helix than the difference with parameter values for another helix.</p></li>\n<li><p>Real tracks are not perfect helix, therefore one has to tweak the parameters to cope with that once one has parameters that are constant for perfect helix.</p></li>\n</ol>",
      "rawMarkdown": "@yuval, I admire your wilI to get others pass you ;)  Let me add two grain of salt to what you shared.\n\n1. We can relax a bit the constraint to have exactly constant parameters when we use DBSCAN, as it can cluster points that are not identical.  It means we don't need to have scanning parameters that yield absolutely constant values for a given helix.  Key is to get parameters closer to each other for a given helix than the difference with parameter values for another helix.\n\n2. Real tracks are not perfect helix, therefore one has to tweak the parameters to cope with that once one has parameters that are constant for perfect helix.",
      "votes": 4,
      "replies": [
        {
          "id": 367748,
          "postDate": "2018-08-08T13:03:02.237Z",
          "content": "<p>@CPMP first I want to congratulate you for the 0.8 (I was waiting for the time you'll be able to pass us).\n I'm and old engineer who is kaggling for fun and I don't give a damn about my LB position. I think the purpose of kaggle is to create knowledge and a lot of work is wasted when people are working on  shaky foundations.</p>\n\n<p>You are correct in both of you observation.  If your features are good and correct you <strong>should be able</strong> to get them to be constant but you <strong>don't have to do it</strong> if your clustering and extending algorithm is good enough. But, I think most of the competitors that are below 0.7 don't have the right features and the clustering algorithm needs to work hard to compensate for it (How much score did you gain whan you got the right accurate features?)</p>\n\n<p><strong>You can't built a tower on shaky foundation</strong></p>",
          "rawMarkdown": "@CPMP first I want to congratulate you for the 0.8 (I was waiting for the time you'll be able to pass us).\n I'm and old engineer who is kaggling for fun and I don't give a damn about my LB position. I think the purpose of kaggle is to create knowledge and a lot of work is wasted when people are working on  shaky foundations.\n\nYou are correct in both of you observation.  If your features are good and correct you **should be able** to get them to be constant but you **don't have to do it** if your clustering and extending algorithm is good enough. But, I think most of the competitors that are below 0.7 don't have the right features and the clustering algorithm needs to work hard to compensate for it (How much score did you gain whan you got the right accurate features?)\n\n**You can't built a tower on shaky foundation**\n\n",
          "votes": 13
        },
        {
          "id": 367750,
          "postDate": "2018-08-08T13:09:38.107Z",
          "content": "<p>Thanks!</p>\n\n<blockquote>\n  <p>How much score did you gain whan you got the right accurate features?</p>\n</blockquote>\n\n<p>Not sure how to answer that, but my progress from 0.70 to 0.785 all comes from better features.</p>",
          "rawMarkdown": "Thanks!\n\n&gt; How much score did you gain whan you got the right accurate features?\n\nNot sure how to answer that, but my progress from 0.70 to 0.785 all comes from better features.",
          "votes": 1
        }
      ]
    },
    {
      "id": 367382,
      "postDate": "2018-08-07T16:06:15.450Z",
      "content": "<p>Thanks for sharing. To check my understanding, if I have features F1 and F2 then: $$ F1_1 = F1_2 = F1_3 = F1_4 = F1_5,  F2_1 = F2_2 = F2_3 = F2_4 = F2_5$$</p>\n\n<p>for 5 hits that are on the path of a perfect helix providing that my parameters {S} are tuned correctly ?</p>",
      "rawMarkdown": "Thanks for sharing. To check my understanding, if I have features F1 and F2 then: $$ F1_1 = F1_2 = F1_3 = F1_4 = F1_5,  F2_1 = F2_2 = F2_3 = F2_4 = F2_5$$\n\nfor 5 hits that are on the path of a perfect helix providing that my parameters {S} are tuned correctly ?",
      "votes": 1,
      "replies": [
        {
          "id": 367397,
          "postDate": "2018-08-07T16:46:53.130Z",
          "content": "<p>You got it right</p>",
          "rawMarkdown": "You got it right",
          "votes": 1
        },
        {
          "id": 367402,
          "postDate": "2018-08-07T17:04:47.807Z",
          "content": "<p>Great! All I need to do now is figure out what these magic features are :)</p>",
          "rawMarkdown": "Great! All I need to do now is figure out what these magic features are :)",
          "votes": 1
        },
        {
          "id": 367865,
          "postDate": "2018-08-08T17:43:23.247Z",
          "content": "<p>Unfortunately, I joined this competition too late. I'm just trying to catch up at this point, I'd have loved to follow up the discussions from the beginning. Right now I'm stuck trying to find the second constant feature. I thought about the slope between the the tangencial movement on the plane XY and Z, supposing that the velocity on Z is constant, but we don't have access to that information I'm afraid. Am I missing something crucial here?</p>",
          "rawMarkdown": "Unfortunately, I joined this competition too late. I'm just trying to catch up at this point, I'd have loved to follow up the discussions from the beginning. Right now I'm stuck trying to find the second constant feature. I thought about the slope between the the tangencial movement on the plane XY and Z, supposing that the velocity on Z is constant, but we don't have access to that information I'm afraid. Am I missing something crucial here?"
        }
      ]
    },
    {
      "id": 367259,
      "postDate": "2018-08-07T11:47:32.427Z",
      "content": "<p>After reading the first paragraph I began to fear that below is a solution. And I can get no silver medal. =)</p>\n\n<p>I will try your suggestion.</p>",
      "rawMarkdown": "After reading the first paragraph I began to fear that below is a solution. And I can get no silver medal. =)\n\nI will try your suggestion.",
      "votes": 2,
      "replies": [
        {
          "id": 367265,
          "postDate": "2018-08-07T12:04:00.690Z",
          "content": "<p>Don't worry I won't give anyone a free ride.</p>",
          "rawMarkdown": "Don't worry I won't give anyone a free ride.",
          "votes": 4
        },
        {
          "id": 367407,
          "postDate": "2018-08-07T17:13:55.990Z",
          "content": "<p>I have checked it. I have good features, but they are sensitive to scanning parameters (S1,S2,..Sm). I think I should take a larger set.</p>",
          "rawMarkdown": "I have checked it. I have good features, but they are sensitive to scanning parameters (S1,S2,..Sm). I think I should take a larger set.",
          "votes": 1
        }
      ]
    },
    {
      "id": 367969,
      "postDate": "2018-08-09T01:04:24.913Z",
      "content": "<p>I've found two features that remain constant in one helix. What next? I've tried to use DBSCAN but it gives 0.0000 for every R, theta0 and z0 I've used. \nI've also tried merge clusters of every DBSCAN, still,​ receive only zeros. </p>",
      "rawMarkdown": "I've found two features that remain constant in one helix. What next? I've tried to use DBSCAN but it gives 0.0000 for every R, theta0 and z0 I've used. \nI've also tried merge clusters of every DBSCAN, still,​ receive only zeros. ",
      "replies": [
        {
          "id": 367984,
          "postDate": "2018-08-09T02:58:06.440Z",
          "content": "<p>Remaining constant for all hits in an helix is  <strong>necessary</strong> but it is not <strong>sufficient</strong>.\nDo your features have different values for different helixes?\nYou can see my post \"Criteria for good features\"  for some other ideas.</p>",
          "rawMarkdown": "Remaining constant for all hits in an helix is  **necessary** but it is not **sufficient**.\nDo your features have different values for different helixes?\nYou can see my post \"Criteria for good features\"  for some other ideas.",
          "votes": 1
        }
      ]
    },
    {
      "id": 367889,
      "postDate": "2018-08-08T18:52:35.397Z",
      "content": "<p>So, I changed features, I think I chose better ones and fewer,  and with weights optimized I got improvement from 0.22 to 0.28 for 5 angles (evenly distributed). I was so happy! ... but when I ran it for 200 angles I got 0.30 instead of 0.5 I had before for \"wrong\" featured.... Too bad! What is wrong?  How many scanning parameters do you use for weights optimization to have adequate indication of the score? I cannot optimize on 1000 values of a scanning parameter due to computational time </p>",
      "rawMarkdown": "So, I changed features, I think I chose better ones and fewer,  and with weights optimized I got improvement from 0.22 to 0.28 for 5 angles (evenly distributed). I was so happy! ... but when I ran it for 200 angles I got 0.30 instead of 0.5 I had before for \"wrong\" featured.... Too bad! What is wrong?  How many scanning parameters do you use for weights optimization to have adequate indication of the score? I cannot optimize on 1000 values of a scanning parameter due to computational time ",
      "replies": [
        {
          "id": 367897,
          "postDate": "2018-08-08T19:16:31.280Z",
          "content": "<p>To get about 0.725 before extension I use about 80000 scanning points (that is including the z shift), but I use binning which is more sensitive. I believe @CPMP is getting there with less then 10000.  </p>",
          "rawMarkdown": "To get about 0.725 before extension I use about 80000 scanning points (that is including the z shift), but I use binning which is more sensitive. I believe @CPMP is getting there with less then 10000.  "
        },
        {
          "id": 367898,
          "postDate": "2018-08-08T19:17:22.787Z",
          "content": "<p>And to get to 0.63 I need only 5500</p>",
          "rawMarkdown": "And to get to 0.63 I need only 5500"
        },
        {
          "id": 367901,
          "postDate": "2018-08-08T19:37:12.360Z",
          "content": "<p>yes, I tried to implement binning from Hough kernel, but it's too complicated for me as i am python beginner (installed python 2 months ago)... my dbscan takes 1 h for 1500 points, so I try to opimise just what I can for 200 points. The question is: when you optimize weights with a naive grid search what the minimum amount you can use before you increase number of points?  I am not going for 0.63 here, i just want to learn as it's my first kaggle try. So, how come I optimised features/weights for 5 angles but got much worse results when I increased number of points? How to optimize then(I presume you do not run grid search on your 80000 points)?</p>",
          "rawMarkdown": "yes, I tried to implement binning from Hough kernel, but it's too complicated for me as i am python beginner (installed python 2 months ago)... my dbscan takes 1 h for 1500 points, so I try to opimise just what I can for 200 points. The question is: when you optimize weights with a naive grid search what the minimum amount you can use before you increase number of points?  I am not going for 0.63 here, i just want to learn as it's my first kaggle try. So, how come I optimised features/weights for 5 angles but got much worse results when I increased number of points? How to optimize then(I presume you do not run grid search on your 80000 points)?"
        },
        {
          "id": 367903,
          "postDate": "2018-08-08T19:42:17.227Z",
          "content": "<p>I thought better features with better weight will get you more even at smaller amount of points, and with number of points increase you just grow your score exponentially and getting to a plato at some stage. But in my case what worked better for 5 points got worse at 200...  I find it weird... </p>\n\n<p>If i need too many points without parallel computing and cloud I would not be able to optimize it. And I am not sure if I succeed with parallel and cloud setup before the deadline (I am very new to that things).</p>\n\n<p>Anyway, I am glad some managed to pass 0.9 target in this competition, I hope it's not a limit and we'll see 0.95 at some stage by the end of the week. And hopefully it will help physicists at LHC to investigate if everything is ok with the Universe on a microlevel or something is wrong :) </p>",
          "rawMarkdown": "I thought better features with better weight will get you more even at smaller amount of points, and with number of points increase you just grow your score exponentially and getting to a plato at some stage. But in my case what worked better for 5 points got worse at 200...  I find it weird... \n\nIf i need too many points without parallel computing and cloud I would not be able to optimize it. And I am not sure if I succeed with parallel and cloud setup before the deadline (I am very new to that things).\n\nAnyway, I am glad some managed to pass 0.9 target in this competition, I hope it's not a limit and we'll see 0.95 at some stage by the end of the week. And hopefully it will help physicists at LHC to investigate if everything is ok with the Universe on a microlevel or something is wrong :) ",
          "votes": 1
        },
        {
          "id": 367906,
          "postDate": "2018-08-08T19:55:18.320Z",
          "content": "<p>&gt; when you optimize weights with a naive grid search what the minimum amount you can use before you increase number of points?</p>\n\n<p>I used 500 points to optimize weights. But not with the grid search. You can use Bayesian optimization.</p>",
          "rawMarkdown": "&gt; when you optimize weights with a naive grid search what the minimum amount you can use before you increase number of points?\n\nI used 500 points to optimize weights. But not with the grid search. You can use Bayesian optimization."
        },
        {
          "id": 367914,
          "postDate": "2018-08-08T20:15:17.090Z",
          "content": "<p>I didn't really optimized weights, just ran the algorithm a few times  adjusted the weight a little, found a good point, and that's it.  My algorithm is not that sensitive to the weights.\nFor the scanning, I randomize  = use random numbers with normal distribution </p>",
          "rawMarkdown": "I didn't really optimized weights, just ran the algorithm a few times  adjusted the weight a little, found a good point, and that's it.  My algorithm is not that sensitive to the weights.\nFor the scanning, I randomize  = use random numbers with normal distribution "
        },
        {
          "id": 368156,
          "postDate": "2018-08-09T11:14:26.907Z",
          "content": "<blockquote>\n  <p>I believe @CPMP is getting there with less then 10000. </p>\n</blockquote>\n\n<p>Right, I get to 0.725 in about 2000 DBSCAN runs now.</p>",
          "rawMarkdown": "&gt; I believe @CPMP is getting there with less then 10000. \n\nRight, I get to 0.725 in about 2000 DBSCAN runs now."
        }
      ]
    },
    {
      "id": 367707,
      "postDate": "2018-08-08T11:06:14.940Z",
      "content": "<p>What the scanning parameters are?</p>",
      "rawMarkdown": "What the scanning parameters are?",
      "replies": [
        {
          "id": 367718,
          "postDate": "2018-08-08T11:48:00.697Z",
          "content": "<p>You need to define them. An example of a common scanning parameter is Z0 (some times called z shifting). </p>",
          "rawMarkdown": "You need to define them. An example of a common scanning parameter is Z0 (some times called z shifting). "
        }
      ]
    },
    {
      "id": 367465,
      "postDate": "2018-08-07T20:44:53.290Z",
      "content": "<p>Thank you! One event takes 30-60 min for me, so I probably would not be able to run the full submission with the new features set one I find then, but I keep going :( That was a useful part of learning! </p>",
      "rawMarkdown": "Thank you! One event takes 30-60 min for me, so I probably would not be able to run the full submission with the new features set one I find then, but I keep going :( That was a useful part of learning! ",
      "replies": [
        {
          "id": 367534,
          "postDate": "2018-08-08T01:15:53.020Z",
          "content": "<p>You can get huge performance improvements by parallelizing the processing of your submission. This package makes it pretty straight forward <a href=\"https://pythonhosted.org/joblib/parallel.html\">https://pythonhosted.org/joblib/parallel.html</a></p>",
          "rawMarkdown": "You can get huge performance improvements by parallelizing the processing of your submission. This package makes it pretty straight forward https://pythonhosted.org/joblib/parallel.html",
          "votes": 1
        },
        {
          "id": 367699,
          "postDate": "2018-08-08T10:40:29.083Z",
          "content": "<p>thank you! more stuff to learn. I am trying to set up pycharm on google cloud to get use of their multiple core CPUs, will more CPU and more RAM be good? Found two good features (but one i used already) !!! I need to get our from this 0.49 score :))) now at 0.51 :)))</p>",
          "rawMarkdown": "thank you! more stuff to learn. I am trying to set up pycharm on google cloud to get use of their multiple core CPUs, will more CPU and more RAM be good? Found two good features (but one i used already) !!! I need to get our from this 0.49 score :))) now at 0.51 :)))"
        },
        {
          "id": 367799,
          "postDate": "2018-08-08T15:14:52.777Z",
          "content": "<p>@Jack Vial thank you Jack, the question is: will it really benefit for core i7, or i need to set up on cloud? I'll do it anyway, it's a nice thing to learn and i started kaggle for learning. It was just not the best choice for the first kaggle, i must admit :), but i learn more then for simpler kaggle playground problem</p>",
          "rawMarkdown": "@Jack Vial thank you Jack, the question is: will it really benefit for core i7, or i need to set up on cloud? I'll do it anyway, it's a nice thing to learn and i started kaggle for learning. It was just not the best choice for the first kaggle, i must admit :), but i learn more then for simpler kaggle playground problem"
        },
        {
          "id": 367813,
          "postDate": "2018-08-08T15:39:20.813Z",
          "content": "<blockquote>\n  <p>will more CPU and more RAM be good? </p>\n</blockquote>\n\n<p>Yes, because you can create submission faster using parallelism, see <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/62883#latest-367811\">https://www.kaggle.com/c/trackml-particle-identification/discussion/62883#latest-367811</a></p>",
          "rawMarkdown": "&gt; will more CPU and more RAM be good? \n\nYes, because you can create submission faster using parallelism, see https://www.kaggle.com/c/trackml-particle-identification/discussion/62883#latest-367811"
        },
        {
          "id": 367972,
          "postDate": "2018-08-09T01:31:42.067Z",
          "content": "<p><a href=\"/blonde\">@blonde</a>, You're welcome. Any multi core processor will benefit from parallelization. Locally I run my validation on a 6 core i5 to process 6 events in parallel. For submissions I use an AWS EC2 instance with more cores.</p>",
          "rawMarkdown": "@blonde, You're welcome. Any multi core processor will benefit from parallelization. Locally I run my validation on a 6 core i5 to process 6 events in parallel. For submissions I use an AWS EC2 instance with more cores."
        }
      ]
    },
    {
      "id": 367832,
      "postDate": "2018-08-08T16:06:40.867Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 367842,
          "postDate": "2018-08-08T16:30:57.337Z",
          "content": "<ol>\n<li><p>I set Theta0 to an arbitrary angle. You could use any number you want</p></li>\n<li><p>To understand X,Y, just set w=2*np.pi and you will get a full circle touching the origin - for r=0 you get X=0 and Y=0</p></li>\n</ol>",
          "rawMarkdown": "1. I set Theta0 to an arbitrary angle. You could use any number you want\n\n2. To understand X,Y, just set w=2*np.pi and you will get a full circle touching the origin - for r=0 you get X=0 and Y=0"
        },
        {
          "id": 367854,
          "postDate": "2018-08-08T17:19:42.297Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 367327,
      "postDate": "2018-08-07T14:08:00.993Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 367730,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-08-08T12:14:45.963000",
      "content": "<p>@yuval, I admire your wilI to get others pass you ;)  Let me add two grain of salt to what you shared.</p>\n\n<ol>\n<li><p>We can relax a bit the constraint to have exactly constant parameters when we use DBSCAN, as it can cluster points that are not identical.  It means we don't need to have scanning parameters that yield absolutely constant values for a given helix.  Key is to get parameters closer to each other for a given helix than the difference with parameter values for another helix.</p></li>\n<li><p>Real tracks are not perfect helix, therefore one has to tweak the parameters to cope with that once one has parameters that are constant for perfect helix.</p></li>\n</ol>",
      "votes": 4,
      "replies": [
        {
          "id": 367748,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-08T13:03:02.237000",
          "content": "<p>@CPMP first I want to congratulate you for the 0.8 (I was waiting for the time you'll be able to pass us).\n I'm and old engineer who is kaggling for fun and I don't give a damn about my LB position. I think the purpose of kaggle is to create knowledge and a lot of work is wasted when people are working on  shaky foundations.</p>\n\n<p>You are correct in both of you observation.  If your features are good and correct you <strong>should be able</strong> to get them to be constant but you <strong>don't have to do it</strong> if your clustering and extending algorithm is good enough. But, I think most of the competitors that are below 0.7 don't have the right features and the clustering algorithm needs to work hard to compensate for it (How much score did you gain whan you got the right accurate features?)</p>\n\n<p><strong>You can't built a tower on shaky foundation</strong></p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 367750,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-08T13:09:38.107000",
          "content": "<p>Thanks!</p>\n\n<blockquote>\n  <p>How much score did you gain whan you got the right accurate features?</p>\n</blockquote>\n\n<p>Not sure how to answer that, but my progress from 0.70 to 0.785 all comes from better features.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 367382,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2018-08-07T16:06:15.450000",
      "content": "<p>Thanks for sharing. To check my understanding, if I have features F1 and F2 then: $$ F1_1 = F1_2 = F1_3 = F1_4 = F1_5,  F2_1 = F2_2 = F2_3 = F2_4 = F2_5$$</p>\n\n<p>for 5 hits that are on the path of a perfect helix providing that my parameters {S} are tuned correctly ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 367397,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-07T16:46:53.130000",
          "content": "<p>You got it right</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367402,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2018-08-07T17:04:47.807000",
          "content": "<p>Great! All I need to do now is figure out what these magic features are :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367865,
          "author_name": "Eduardo Konishi",
          "author_url": "",
          "post_date": "2018-08-08T17:43:23.247000",
          "content": "<p>Unfortunately, I joined this competition too late. I'm just trying to catch up at this point, I'd have loved to follow up the discussions from the beginning. Right now I'm stuck trying to find the second constant feature. I thought about the slope between the the tangencial movement on the plane XY and Z, supposing that the velocity on Z is constant, but we don't have access to that information I'm afraid. Am I missing something crucial here?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 367259,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2018-08-07T11:47:32.427000",
      "content": "<p>After reading the first paragraph I began to fear that below is a solution. And I can get no silver medal. =)</p>\n\n<p>I will try your suggestion.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 367265,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-07T12:04:00.690000",
          "content": "<p>Don't worry I won't give anyone a free ride.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 367407,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-07T17:13:55.990000",
          "content": "<p>I have checked it. I have good features, but they are sensitive to scanning parameters (S1,S2,..Sm). I think I should take a larger set.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 367969,
      "author_name": "Daniil  Okhlopkov",
      "author_url": "",
      "post_date": "2018-08-09T01:04:24.913000",
      "content": "<p>I've found two features that remain constant in one helix. What next? I've tried to use DBSCAN but it gives 0.0000 for every R, theta0 and z0 I've used. \nI've also tried merge clusters of every DBSCAN, still,​ receive only zeros. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 367984,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-09T02:58:06.440000",
          "content": "<p>Remaining constant for all hits in an helix is  <strong>necessary</strong> but it is not <strong>sufficient</strong>.\nDo your features have different values for different helixes?\nYou can see my post \"Criteria for good features\"  for some other ideas.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 367889,
      "author_name": "Blonde",
      "author_url": "",
      "post_date": "2018-08-08T18:52:35.397000",
      "content": "<p>So, I changed features, I think I chose better ones and fewer,  and with weights optimized I got improvement from 0.22 to 0.28 for 5 angles (evenly distributed). I was so happy! ... but when I ran it for 200 angles I got 0.30 instead of 0.5 I had before for \"wrong\" featured.... Too bad! What is wrong?  How many scanning parameters do you use for weights optimization to have adequate indication of the score? I cannot optimize on 1000 values of a scanning parameter due to computational time </p>",
      "votes": 0,
      "replies": [
        {
          "id": 367897,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-08T19:16:31.280000",
          "content": "<p>To get about 0.725 before extension I use about 80000 scanning points (that is including the z shift), but I use binning which is more sensitive. I believe @CPMP is getting there with less then 10000.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367898,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-08T19:17:22.787000",
          "content": "<p>And to get to 0.63 I need only 5500</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367901,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-08-08T19:37:12.360000",
          "content": "<p>yes, I tried to implement binning from Hough kernel, but it's too complicated for me as i am python beginner (installed python 2 months ago)... my dbscan takes 1 h for 1500 points, so I try to opimise just what I can for 200 points. The question is: when you optimize weights with a naive grid search what the minimum amount you can use before you increase number of points?  I am not going for 0.63 here, i just want to learn as it's my first kaggle try. So, how come I optimised features/weights for 5 angles but got much worse results when I increased number of points? How to optimize then(I presume you do not run grid search on your 80000 points)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367903,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-08-08T19:42:17.227000",
          "content": "<p>I thought better features with better weight will get you more even at smaller amount of points, and with number of points increase you just grow your score exponentially and getting to a plato at some stage. But in my case what worked better for 5 points got worse at 200...  I find it weird... </p>\n\n<p>If i need too many points without parallel computing and cloud I would not be able to optimize it. And I am not sure if I succeed with parallel and cloud setup before the deadline (I am very new to that things).</p>\n\n<p>Anyway, I am glad some managed to pass 0.9 target in this competition, I hope it's not a limit and we'll see 0.95 at some stage by the end of the week. And hopefully it will help physicists at LHC to investigate if everything is ok with the Universe on a microlevel or something is wrong :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367906,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-08T19:55:18.320000",
          "content": "<p>&gt; when you optimize weights with a naive grid search what the minimum amount you can use before you increase number of points?</p>\n\n<p>I used 500 points to optimize weights. But not with the grid search. You can use Bayesian optimization.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367914,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-08T20:15:17.090000",
          "content": "<p>I didn't really optimized weights, just ran the algorithm a few times  adjusted the weight a little, found a good point, and that's it.  My algorithm is not that sensitive to the weights.\nFor the scanning, I randomize  = use random numbers with normal distribution </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 368156,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-09T11:14:26.907000",
          "content": "<blockquote>\n  <p>I believe @CPMP is getting there with less then 10000. </p>\n</blockquote>\n\n<p>Right, I get to 0.725 in about 2000 DBSCAN runs now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 367707,
      "author_name": "Konstantin Gavrilchik",
      "author_url": "",
      "post_date": "2018-08-08T11:06:14.940000",
      "content": "<p>What the scanning parameters are?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 367718,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-08T11:48:00.697000",
          "content": "<p>You need to define them. An example of a common scanning parameter is Z0 (some times called z shifting). </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 367465,
      "author_name": "Blonde",
      "author_url": "",
      "post_date": "2018-08-07T20:44:53.290000",
      "content": "<p>Thank you! One event takes 30-60 min for me, so I probably would not be able to run the full submission with the new features set one I find then, but I keep going :( That was a useful part of learning! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 367534,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2018-08-08T01:15:53.020000",
          "content": "<p>You can get huge performance improvements by parallelizing the processing of your submission. This package makes it pretty straight forward <a href=\"https://pythonhosted.org/joblib/parallel.html\">https://pythonhosted.org/joblib/parallel.html</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367699,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-08-08T10:40:29.083000",
          "content": "<p>thank you! more stuff to learn. I am trying to set up pycharm on google cloud to get use of their multiple core CPUs, will more CPU and more RAM be good? Found two good features (but one i used already) !!! I need to get our from this 0.49 score :))) now at 0.51 :)))</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367799,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-08-08T15:14:52.777000",
          "content": "<p>@Jack Vial thank you Jack, the question is: will it really benefit for core i7, or i need to set up on cloud? I'll do it anyway, it's a nice thing to learn and i started kaggle for learning. It was just not the best choice for the first kaggle, i must admit :), but i learn more then for simpler kaggle playground problem</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367813,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-08T15:39:20.813000",
          "content": "<blockquote>\n  <p>will more CPU and more RAM be good? </p>\n</blockquote>\n\n<p>Yes, because you can create submission faster using parallelism, see <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/62883#latest-367811\">https://www.kaggle.com/c/trackml-particle-identification/discussion/62883#latest-367811</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367972,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2018-08-09T01:31:42.067000",
          "content": "<p><a href=\"/blonde\">@blonde</a>, You're welcome. Any multi core processor will benefit from parallelization. Locally I run my validation on a 6 core i5 to process 6 events in parallel. For submissions I use an AWS EC2 instance with more cores.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 367832,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-08-08T16:06:40.867000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 367842,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2018-08-08T16:30:57.337000",
          "content": "<ol>\n<li><p>I set Theta0 to an arbitrary angle. You could use any number you want</p></li>\n<li><p>To understand X,Y, just set w=2*np.pi and you will get a full circle touching the origin - for r=0 you get X=0 and Y=0</p></li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 367854,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-08-08T17:19:42.297000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 367327,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-08-07T14:08:00.993000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "367235": "This post is written  for kagglers who are using clustering as their main track detection mechanism and are still below LB of 0.7.\n\nIf this is your situation,  probably you have an issue with your features.\nTo understand and fix this issue, do the following:\n\n1. Run this code:\n\nR=2000    #Helix Radius\n\nw=np.pi/3 #length of arc\n\nv=3000    # velocity in Z\n\ntheta0=-0.3  # angle at the origin\n\nD=0    #D=500   -  Use for helixes far from the origin\n\nz0=0    #z0=200   -  Z shifting\n\nt=np.random.rand(50)\n\nX=-R*np.sin(w*t+theta0)+(R-D)*np.sin(theta0)\n\nY=-R*np.cos(w*t+theta0)+(R-D)*np.cos(theta0)\n\nZ=z0+v*t\n\n\nNow you have a simulation of a perfect helix, starting at the origin (X,Y,Z) \n\n2. calculate your features for this helix\n\n3. if your features are (F1, F2, ...Fn)  and your scanning parameters are (S1,S2,..Sm), you should be able to find a set of {S} for which {F} are constant for all values of `t` \n\nBe ware \"**constant**\" is **constant**  not approximately the same. \nIf you can't get the features to be constant you don't have the right features\n(use as few features and scanning parameters as you can - less is more) \n",
    "367730": "@yuval, I admire your wilI to get others pass you ;)  Let me add two grain of salt to what you shared.\n\n1. We can relax a bit the constraint to have exactly constant parameters when we use DBSCAN, as it can cluster points that are not identical.  It means we don't need to have scanning parameters that yield absolutely constant values for a given helix.  Key is to get parameters closer to each other for a given helix than the difference with parameter values for another helix.\n\n2. Real tracks are not perfect helix, therefore one has to tweak the parameters to cope with that once one has parameters that are constant for perfect helix.",
    "367382": "Thanks for sharing. To check my understanding, if I have features F1 and F2 then: $$ F1_1 = F1_2 = F1_3 = F1_4 = F1_5,  F2_1 = F2_2 = F2_3 = F2_4 = F2_5$$\n\nfor 5 hits that are on the path of a perfect helix providing that my parameters {S} are tuned correctly ?",
    "367259": "After reading the first paragraph I began to fear that below is a solution. And I can get no silver medal. =)\n\nI will try your suggestion.",
    "367969": "I've found two features that remain constant in one helix. What next? I've tried to use DBSCAN but it gives 0.0000 for every R, theta0 and z0 I've used. \nI've also tried merge clusters of every DBSCAN, still,​ receive only zeros. ",
    "367889": "So, I changed features, I think I chose better ones and fewer,  and with weights optimized I got improvement from 0.22 to 0.28 for 5 angles (evenly distributed). I was so happy! ... but when I ran it for 200 angles I got 0.30 instead of 0.5 I had before for \"wrong\" featured.... Too bad! What is wrong?  How many scanning parameters do you use for weights optimization to have adequate indication of the score? I cannot optimize on 1000 values of a scanning parameter due to computational time ",
    "367707": "What the scanning parameters are?",
    "367465": "Thank you! One event takes 30-60 min for me, so I probably would not be able to run the full submission with the new features set one I find then, but I keep going :( That was a useful part of learning! ",
    "367832": "",
    "367327": ""
  }
}