{
  "id": 58323,
  "title": "ensemble clustering?",
  "url": "/competitions/trackml-particle-identification/discussion/58323",
  "author_name": "hengck23",
  "post_date": "2018-06-06T05:11:58.288000",
  "votes": 5,
  "comment_count": 16,
  "views": 0,
  "content": "<p>anyone has good suggestion how to ensemble results for track reconstruction? </p>",
  "messages": [
    {
      "id": 338970,
      "postDate": "2018-06-06T05:11:58.290Z",
      "content": "<p>anyone has good suggestion how to ensemble results for track reconstruction? </p>",
      "rawMarkdown": "anyone has good suggestion how to ensemble results for track reconstruction? ",
      "votes": 5
    },
    {
      "id": 367711,
      "postDate": "2018-08-08T11:12:51.483Z",
      "content": "<p>Here is a way to use ML to do the clustering:</p>\n\n<ol>\n<li><p>write a new score function to measure the purity of a given candidate track:\n   e.g. purity( point1,point2,point3...) = num_of_point_with_same_non_zero_particle_id / total_num_of_point</p></li>\n<li><p>run DBSCAN. Use different helix unrolling parameters. Adjust EPS so that purity is high (preferably 1). It is ok to break long tracks into smaller ones</p></li>\n<li><p>for each DBSCAN results, for each point construct:\n  feature = length_of_cluster, x,y,x, .... etc ... (other example feature can be x,yz, of closest point, layer_id, direction vector to next/prev layer point ...)</p></li>\n<li><p>Now for each point, you have:\n    f =  [ feature1 from DBSCAN1, feature2 from DBSCAN2, .... featureN from DBSCANN ]</p>\n\n<p>the concatenation of N DBSCAN results capture the neighbor characteristics of a point.</p></li>\n<li><p>Train a classifier to map f to y= [ 1, 0,0, ....1, 0 ] (e.g. deep learning or random trees?)\ndim of y is N. \"1\" at n-th indicate DBSCAN-N results is pure. Use sigmoid because several DBSCAN can be correct</p></li>\n<li><p>Now you know which DBSCAN results is correct. you can construct adjacency grpah or use other methods to fuse small tracks into large ones.</p></li>\n</ol>",
      "rawMarkdown": "Here is a way to use ML to do the clustering:\n\n1. write a new score function to measure the purity of a given candidate track:\n       e.g. purity( point1,point2,point3...) = num_of_point_with_same_non_zero_particle_id / total_num_of_point\n\n2. run DBSCAN. Use different helix unrolling parameters. Adjust EPS so that purity is high (preferably 1). It is ok to break long tracks into smaller ones\n\n3. for each DBSCAN results, for each point construct:\n      feature = length_of_cluster, x,y,x, .... etc ... (other example feature can be x,yz, of closest point, layer_id, direction vector to next/prev layer point ...)\n    \n4. Now for each point, you have:\n        f =  [ feature1 from DBSCAN1, feature2 from DBSCAN2, .... featureN from DBSCANN ]\n\n       the concatenation of N DBSCAN results capture the neighbor characteristics of a point.\n\n5. Train a classifier to map f to y= [ 1, 0,0, ....1, 0 ] (e.g. deep learning or random trees?)\n    dim of y is N. \"1\" at n-th indicate DBSCAN-N results is pure. Use sigmoid because several DBSCAN can be correct\n\n6. Now you know which DBSCAN results is correct. you can construct adjacency grpah or use other methods to fuse small tracks into large ones.\n\n\n\n \n\n",
      "votes": 1,
      "replies": [
        {
          "id": 367911,
          "postDate": "2018-08-08T20:03:03.480Z",
          "content": "<p>Thank you Heng. I just realized an obvious thing: to do ensembling you need to use models with different features, not just models with very similar features and twisted around parameters.... so more work on features for me.</p>\n\n<p>When this competition is finished maybe it's worth creating github library of your good codes so people can do similar things it future (like ensembling of clusters, just as idea). It may help Big Theoretical Physics as well and hence our understanding of the Universe may gain new knowledge... sometimes little things contribute to great impacts</p>",
          "rawMarkdown": "Thank you Heng. I just realized an obvious thing: to do ensembling you need to use models with different features, not just models with very similar features and twisted around parameters.... so more work on features for me.\n\nWhen this competition is finished maybe it's worth creating github library of your good codes so people can do similar things it future (like ensembling of clusters, just as idea). It may help Big Theoretical Physics as well and hence our understanding of the Universe may gain new knowledge... sometimes little things contribute to great impacts"
        }
      ]
    },
    {
      "id": 367475,
      "postDate": "2018-08-07T21:18:19.263Z",
      "content": "<p>Dear @Heng Thanks for this discussion! And for links. I am the first time kaggler, learning how to ensemble. Yes, this competition is not the best for the start, but after all this time I invested in it I have to ensemble those models !!! You and @CPMP are here for long, maybe you know if there were discussion and kernels in the previous clustering competitions on the models ensembling for clustering like that? I have a few models, all relatively low scores, and want to ensemble them  to 1. learn how to ensemble and\n 2. to improve my miserable score :)</p>\n\n<p>I found weighted blend by @Sergey Zlobin, and few more <a href=\"https://www.kaggle.com/sergeyzlobin/wanna-blend\">https://www.kaggle.com/sergeyzlobin/wanna-blend</a> and here \nand some more here <a href=\"https://www.kaggle.com/vincento/ensemble-your-models-for-better-score\">https://www.kaggle.com/vincento/ensemble-your-models-for-better-score</a>\nI found this one useful\n<a href=\"https://www.kaggle.com/derrickchua29/ensemble-models-comparison-techniques\">https://www.kaggle.com/derrickchua29/ensemble-models-comparison-techniques</a></p>\n\n<p>I am a first time kaggler with no background in ML/programming, so forgive me if I am wrong. Non of them is really appropriate for this competition and here we need to choose the longest track or so. We can use group by and compare the tracks in different models, and then choose the best one... but how do we make sure we do not assign one hit to two different tracks? How to make sure all hits are also assigned to some tracks? create track 0 where all undefined stuff goes?  I started this competition with googling \"how to install python\" and \"how to use github\",  totally new... sorry if this question is inappropriate, but maybe there is a good material for beginners that i missed to make through in this ensembling for this unusual clustering thing... I believe people with silver-gold medal results know how to ensemble, so only novice people like me struggle with this problem. So, if you could give us few clarifications it would not put you at risk :)  Yes, i saw those python libruary, but it's plug and play, and I need to learn how to do this to be able to do ensembling in future</p>",
      "rawMarkdown": "Dear @Heng Thanks for this discussion! And for links. I am the first time kaggler, learning how to ensemble. Yes, this competition is not the best for the start, but after all this time I invested in it I have to ensemble those models !!! You and @CPMP are here for long, maybe you know if there were discussion and kernels in the previous clustering competitions on the models ensembling for clustering like that? I have a few models, all relatively low scores, and want to ensemble them  to 1. learn how to ensemble and\n 2. to improve my miserable score :)\n\nI found weighted blend by @Sergey Zlobin, and few more https://www.kaggle.com/sergeyzlobin/wanna-blend and here \nand some more here https://www.kaggle.com/vincento/ensemble-your-models-for-better-score\nI found this one useful\nhttps://www.kaggle.com/derrickchua29/ensemble-models-comparison-techniques\n\nI am a first time kaggler with no background in ML/programming, so forgive me if I am wrong. Non of them is really appropriate for this competition and here we need to choose the longest track or so. We can use group by and compare the tracks in different models, and then choose the best one... but how do we make sure we do not assign one hit to two different tracks? How to make sure all hits are also assigned to some tracks? create track 0 where all undefined stuff goes?  I started this competition with googling \"how to install python\" and \"how to use github\",  totally new... sorry if this question is inappropriate, but maybe there is a good material for beginners that i missed to make through in this ensembling for this unusual clustering thing... I believe people with silver-gold medal results know how to ensemble, so only novice people like me struggle with this problem. So, if you could give us few clarifications it would not put you at risk :)  Yes, i saw those python libruary, but it's plug and play, and I need to learn how to do this to be able to do ensembling in future",
      "votes": 1,
      "replies": [
        {
          "id": 367580,
          "postDate": "2018-08-08T03:53:20.570Z",
          "content": "<p>Wow, you're courageous to start your first competition with this one.  This one is very unusual and very challenging.  If you really are a beginner in ML then I would encourage you to look at the playground competitions first.</p>\n\n<p>To your question, I have not seen ensembling of clusters in past competitions, this is a new topic for me (and for many of us I guess).  Heng shared several relevant work, but I think what we have is very specific, and specific methods based on track length and other properties can be way more effective than general purpose method.</p>\n\n<p>And to avoid assign a hit to more than one track is very simple: just have one feature per hit called 'track_id'.  Use the 0 value as a catch all for all hits not assigned to a track.  </p>",
          "rawMarkdown": "Wow, you're courageous to start your first competition with this one.  This one is very unusual and very challenging.  If you really are a beginner in ML then I would encourage you to look at the playground competitions first.\n\nTo your question, I have not seen ensembling of clusters in past competitions, this is a new topic for me (and for many of us I guess).  Heng shared several relevant work, but I think what we have is very specific, and specific methods based on track length and other properties can be way more effective than general purpose method.\n\nAnd to avoid assign a hit to more than one track is very simple: just have one feature per hit called 'track_id'.  Use the 0 value as a catch all for all hits not assigned to a track.  ",
          "votes": 2
        },
        {
          "id": 367705,
          "postDate": "2018-08-08T10:55:58.920Z",
          "content": "<p>@CPMP I just love physics (I did PhD in physics before this maternity, but in experimental physics, lasers, nothing in common with particles, so it did not help me at all :)) And yes, that was a bad idea to start with this one, i should have started with titanic :), and it was a bad idea to try to implement unet from Heng and LSTM from other papers on this LHC problem with my zero background in deep learning, I just killed a lot of time... And I should have read the discussions 2 months ago instead of trying to figure things out myself ignoring discussion session at all... I  read it only last week and regret I did not do it before. Anyway, kaggle was also new for me, lot's of things to learn, next competition will do better. </p>\n\n<p>Yes, it's very non-standard ensembling, I noticed. I also wanted to choose the longest tracks, my zero python background does not help with implementation :))) I keep posting my weak tries in kernels and hopefully some other beginners will help and maybe we all get through...</p>",
          "rawMarkdown": "@CPMP I just love physics (I did PhD in physics before this maternity, but in experimental physics, lasers, nothing in common with particles, so it did not help me at all :)) And yes, that was a bad idea to start with this one, i should have started with titanic :), and it was a bad idea to try to implement unet from Heng and LSTM from other papers on this LHC problem with my zero background in deep learning, I just killed a lot of time... And I should have read the discussions 2 months ago instead of trying to figure things out myself ignoring discussion session at all... I  read it only last week and regret I did not do it before. Anyway, kaggle was also new for me, lot's of things to learn, next competition will do better. \n\nYes, it's very non-standard ensembling, I noticed. I also wanted to choose the longest tracks, my zero python background does not help with implementation :))) I keep posting my weak tries in kernels and hopefully some other beginners will help and maybe we all get through...",
          "votes": 2
        }
      ]
    },
    {
      "id": 343864,
      "postDate": "2018-06-16T11:34:20.413Z",
      "content": "<p>Ensembling 5 models gave us a 0.05 boost but then we hit the wall. It seems no matter what the 6th model is, it won't give us more benefit. From our experience, ensembling models with high quality tracks gives us more accuracy than ensembling strong models + weak models. We tried to ensemble models with a cone slicing model, but it degrades after the 5th model, so we had to remove it. Our ensembling code still needs to be improved since @Grzegorz proved it gave him a higher bump.</p>",
      "rawMarkdown": "Ensembling 5 models gave us a 0.05 boost but then we hit the wall. It seems no matter what the 6th model is, it won't give us more benefit. From our experience, ensembling models with high quality tracks gives us more accuracy than ensembling strong models + weak models. We tried to ensemble models with a cone slicing model, but it degrades after the 5th model, so we had to remove it. Our ensembling code still needs to be improved since @Grzegorz proved it gave him a higher bump.\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 343890,
          "postDate": "2018-06-16T12:35:03.850Z",
          "content": "<p>@Nicole, I have no IT background, so it is easier to me to invent a new ensembling method dedicated for this problem and code it from scratch than to find, understand and implement/modify commonly used general one.</p>",
          "rawMarkdown": "@Nicole, I have no IT background, so it is easier to me to invent a new ensembling method dedicated for this problem and code it from scratch than to find, understand and implement/modify commonly used general one."
        },
        {
          "id": 343908,
          "postDate": "2018-06-16T13:52:56.810Z",
          "content": "<p>@Grezgorz, I have no physics/chemistry/ML background, so it's easier for me to copy your formulae :D and use the rookie approach - trial and error + naive grid search instead of looking-very-cool-Bayesian Optimization to find the optimal parameters. ;) </p>",
          "rawMarkdown": "@Grezgorz, I have no physics/chemistry/ML background, so it's easier for me to copy your formulae :D and use the rookie approach - trial and error + naive grid search instead of looking-very-cool-Bayesian Optimization to find the optimal parameters. ;) \n"
        }
      ]
    },
    {
      "id": 340120,
      "postDate": "2018-06-08T12:00:20.553Z",
      "content": "<p>I use an ensembler integrated with the models. Ensembling is partially performed on the fly (while calculations are not finished yet) changing a little bit the state of each model calculations giving them a chance for additional improvement. The boost is ~0.08 for 7 models.</p>",
      "rawMarkdown": "I use an ensembler integrated with the models. Ensembling is partially performed on the fly (while calculations are not finished yet) changing a little bit the state of each model calculations giving them a chance for additional improvement. The boost is ~0.08 for 7 models.\n",
      "votes": 1
    },
    {
      "id": 339144,
      "postDate": "2018-06-06T11:50:59.050Z",
      "content": "<p>I use largest track to merge, and it seems this is also used in some of the public kernels.</p>",
      "rawMarkdown": "I use largest track to merge, and it seems this is also used in some of the public kernels.",
      "votes": 1
    },
    {
      "id": 340141,
      "postDate": "2018-06-08T13:00:26.553Z",
      "content": "<p>some notes on \"ensemble clustering\"</p>\n\n<p>@CPMP python scoring function is useful for relabeling in majority voting.</p>\n\n<p><a href=\"https://www.kaggle.com/cpmpml/a-faster-python-scoring-function\">https://www.kaggle.com/cpmpml/a-faster-python-scoring-function</a></p>\n\n<p><a href=\"https://cse.buffalo.edu/~jing/cse601/fa12/materials/clustering_ensemble.pdf\">https://cse.buffalo.edu/~jing/cse601/fa12/materials/clustering_ensemble.pdf</a></p>\n\n<p>see also:</p>\n\n<p><a href=\"https://datascience.stackexchange.com/questions/27420/how-to-apply-ensemble-clustering-method\">https://datascience.stackexchange.com/questions/27420/how-to-apply-ensemble-clustering-method</a></p>",
      "rawMarkdown": "some notes on \"ensemble clustering\"\n\n@CPMP python scoring function is useful for relabeling in majority voting.\n\nhttps://www.kaggle.com/cpmpml/a-faster-python-scoring-function\n\nhttps://cse.buffalo.edu/~jing/cse601/fa12/materials/clustering_ensemble.pdf\n\nsee also:\n\nhttps://datascience.stackexchange.com/questions/27420/how-to-apply-ensemble-clustering-method\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 340220,
          "postDate": "2018-06-08T16:43:15.243Z",
          "content": "<p>Thanks, (and not just for citing me)!  Have you tried the Python library?  I probably will.</p>",
          "rawMarkdown": "Thanks, (and not just for citing me)!  Have you tried the Python library?  I probably will."
        },
        {
          "id": 366238,
          "postDate": "2018-08-04T13:19:36.893Z",
          "content": "<blockquote>\n  <p>Have you tried the Python library?</p>\n</blockquote>\n\n<p>I have tried now, but it doesn't support Windows. :(\nI have not found another implementations for Python.</p>",
          "rawMarkdown": "&gt; Have you tried the Python library?\n\nI have tried now, but it doesn't support Windows. :(\nI have not found another implementations for Python."
        },
        {
          "id": 366543,
          "postDate": "2018-08-05T18:31:44.070Z",
          "content": "<p>I tried the library on Mac. It took a little bit of torture to use the library.\nUnfortunately I didn't get a result in several hours even for one event. So I drop it.\nHowever I don't wait good results from a general algorithm. </p>",
          "rawMarkdown": "I tried the library on Mac. It took a little bit of torture to use the library.\nUnfortunately I didn't get a result in several hours even for one event. So I drop it.\nHowever I don't wait good results from a general algorithm. "
        }
      ]
    },
    {
      "id": 340121,
      "postDate": "2018-06-08T12:01:45.470Z",
      "content": "<p>@Grzegorz Sionkowski</p>\n\n<p>Thanks for the information! </p>",
      "rawMarkdown": "@Grzegorz Sionkowski\n \nThanks for the information! \n"
    },
    {
      "id": 367909,
      "postDate": "2018-08-08T20:01:54.510Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 367711,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-08-08T11:12:51.483000",
      "content": "<p>Here is a way to use ML to do the clustering:</p>\n\n<ol>\n<li><p>write a new score function to measure the purity of a given candidate track:\n   e.g. purity( point1,point2,point3...) = num_of_point_with_same_non_zero_particle_id / total_num_of_point</p></li>\n<li><p>run DBSCAN. Use different helix unrolling parameters. Adjust EPS so that purity is high (preferably 1). It is ok to break long tracks into smaller ones</p></li>\n<li><p>for each DBSCAN results, for each point construct:\n  feature = length_of_cluster, x,y,x, .... etc ... (other example feature can be x,yz, of closest point, layer_id, direction vector to next/prev layer point ...)</p></li>\n<li><p>Now for each point, you have:\n    f =  [ feature1 from DBSCAN1, feature2 from DBSCAN2, .... featureN from DBSCANN ]</p>\n\n<p>the concatenation of N DBSCAN results capture the neighbor characteristics of a point.</p></li>\n<li><p>Train a classifier to map f to y= [ 1, 0,0, ....1, 0 ] (e.g. deep learning or random trees?)\ndim of y is N. \"1\" at n-th indicate DBSCAN-N results is pure. Use sigmoid because several DBSCAN can be correct</p></li>\n<li><p>Now you know which DBSCAN results is correct. you can construct adjacency grpah or use other methods to fuse small tracks into large ones.</p></li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 367911,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-08-08T20:03:03.480000",
          "content": "<p>Thank you Heng. I just realized an obvious thing: to do ensembling you need to use models with different features, not just models with very similar features and twisted around parameters.... so more work on features for me.</p>\n\n<p>When this competition is finished maybe it's worth creating github library of your good codes so people can do similar things it future (like ensembling of clusters, just as idea). It may help Big Theoretical Physics as well and hence our understanding of the Universe may gain new knowledge... sometimes little things contribute to great impacts</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 367475,
      "author_name": "Blonde",
      "author_url": "",
      "post_date": "2018-08-07T21:18:19.263000",
      "content": "<p>Dear @Heng Thanks for this discussion! And for links. I am the first time kaggler, learning how to ensemble. Yes, this competition is not the best for the start, but after all this time I invested in it I have to ensemble those models !!! You and @CPMP are here for long, maybe you know if there were discussion and kernels in the previous clustering competitions on the models ensembling for clustering like that? I have a few models, all relatively low scores, and want to ensemble them  to 1. learn how to ensemble and\n 2. to improve my miserable score :)</p>\n\n<p>I found weighted blend by @Sergey Zlobin, and few more <a href=\"https://www.kaggle.com/sergeyzlobin/wanna-blend\">https://www.kaggle.com/sergeyzlobin/wanna-blend</a> and here \nand some more here <a href=\"https://www.kaggle.com/vincento/ensemble-your-models-for-better-score\">https://www.kaggle.com/vincento/ensemble-your-models-for-better-score</a>\nI found this one useful\n<a href=\"https://www.kaggle.com/derrickchua29/ensemble-models-comparison-techniques\">https://www.kaggle.com/derrickchua29/ensemble-models-comparison-techniques</a></p>\n\n<p>I am a first time kaggler with no background in ML/programming, so forgive me if I am wrong. Non of them is really appropriate for this competition and here we need to choose the longest track or so. We can use group by and compare the tracks in different models, and then choose the best one... but how do we make sure we do not assign one hit to two different tracks? How to make sure all hits are also assigned to some tracks? create track 0 where all undefined stuff goes?  I started this competition with googling \"how to install python\" and \"how to use github\",  totally new... sorry if this question is inappropriate, but maybe there is a good material for beginners that i missed to make through in this ensembling for this unusual clustering thing... I believe people with silver-gold medal results know how to ensemble, so only novice people like me struggle with this problem. So, if you could give us few clarifications it would not put you at risk :)  Yes, i saw those python libruary, but it's plug and play, and I need to learn how to do this to be able to do ensembling in future</p>",
      "votes": 1,
      "replies": [
        {
          "id": 367580,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-08-08T03:53:20.570000",
          "content": "<p>Wow, you're courageous to start your first competition with this one.  This one is very unusual and very challenging.  If you really are a beginner in ML then I would encourage you to look at the playground competitions first.</p>\n\n<p>To your question, I have not seen ensembling of clusters in past competitions, this is a new topic for me (and for many of us I guess).  Heng shared several relevant work, but I think what we have is very specific, and specific methods based on track length and other properties can be way more effective than general purpose method.</p>\n\n<p>And to avoid assign a hit to more than one track is very simple: just have one feature per hit called 'track_id'.  Use the 0 value as a catch all for all hits not assigned to a track.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 367705,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-08-08T10:55:58.920000",
          "content": "<p>@CPMP I just love physics (I did PhD in physics before this maternity, but in experimental physics, lasers, nothing in common with particles, so it did not help me at all :)) And yes, that was a bad idea to start with this one, i should have started with titanic :), and it was a bad idea to try to implement unet from Heng and LSTM from other papers on this LHC problem with my zero background in deep learning, I just killed a lot of time... And I should have read the discussions 2 months ago instead of trying to figure things out myself ignoring discussion session at all... I  read it only last week and regret I did not do it before. Anyway, kaggle was also new for me, lot's of things to learn, next competition will do better. </p>\n\n<p>Yes, it's very non-standard ensembling, I noticed. I also wanted to choose the longest tracks, my zero python background does not help with implementation :))) I keep posting my weak tries in kernels and hopefully some other beginners will help and maybe we all get through...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 343864,
      "author_name": "Nicole Finnie",
      "author_url": "",
      "post_date": "2018-06-16T11:34:20.413000",
      "content": "<p>Ensembling 5 models gave us a 0.05 boost but then we hit the wall. It seems no matter what the 6th model is, it won't give us more benefit. From our experience, ensembling models with high quality tracks gives us more accuracy than ensembling strong models + weak models. We tried to ensemble models with a cone slicing model, but it degrades after the 5th model, so we had to remove it. Our ensembling code still needs to be improved since @Grzegorz proved it gave him a higher bump.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 343890,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-06-16T12:35:03.850000",
          "content": "<p>@Nicole, I have no IT background, so it is easier to me to invent a new ensembling method dedicated for this problem and code it from scratch than to find, understand and implement/modify commonly used general one.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343908,
          "author_name": "Nicole Finnie",
          "author_url": "",
          "post_date": "2018-06-16T13:52:56.810000",
          "content": "<p>@Grezgorz, I have no physics/chemistry/ML background, so it's easier for me to copy your formulae :D and use the rookie approach - trial and error + naive grid search instead of looking-very-cool-Bayesian Optimization to find the optimal parameters. ;) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 340120,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2018-06-08T12:00:20.553000",
      "content": "<p>I use an ensembler integrated with the models. Ensembling is partially performed on the fly (while calculations are not finished yet) changing a little bit the state of each model calculations giving them a chance for additional improvement. The boost is ~0.08 for 7 models.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 339144,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-06-06T11:50:59.050000",
      "content": "<p>I use largest track to merge, and it seems this is also used in some of the public kernels.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 340141,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-06-08T13:00:26.553000",
      "content": "<p>some notes on \"ensemble clustering\"</p>\n\n<p>@CPMP python scoring function is useful for relabeling in majority voting.</p>\n\n<p><a href=\"https://www.kaggle.com/cpmpml/a-faster-python-scoring-function\">https://www.kaggle.com/cpmpml/a-faster-python-scoring-function</a></p>\n\n<p><a href=\"https://cse.buffalo.edu/~jing/cse601/fa12/materials/clustering_ensemble.pdf\">https://cse.buffalo.edu/~jing/cse601/fa12/materials/clustering_ensemble.pdf</a></p>\n\n<p>see also:</p>\n\n<p><a href=\"https://datascience.stackexchange.com/questions/27420/how-to-apply-ensemble-clustering-method\">https://datascience.stackexchange.com/questions/27420/how-to-apply-ensemble-clustering-method</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 340220,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-08T16:43:15.243000",
          "content": "<p>Thanks, (and not just for citing me)!  Have you tried the Python library?  I probably will.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366238,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-04T13:19:36.893000",
          "content": "<blockquote>\n  <p>Have you tried the Python library?</p>\n</blockquote>\n\n<p>I have tried now, but it doesn't support Windows. :(\nI have not found another implementations for Python.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 366543,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2018-08-05T18:31:44.070000",
          "content": "<p>I tried the library on Mac. It took a little bit of torture to use the library.\nUnfortunately I didn't get a result in several hours even for one event. So I drop it.\nHowever I don't wait good results from a general algorithm. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 340121,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-06-08T12:01:45.470000",
      "content": "<p>@Grzegorz Sionkowski</p>\n\n<p>Thanks for the information! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 367909,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-08-08T20:01:54.510000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "338970": "anyone has good suggestion how to ensemble results for track reconstruction? ",
    "367711": "Here is a way to use ML to do the clustering:\n\n1. write a new score function to measure the purity of a given candidate track:\n       e.g. purity( point1,point2,point3...) = num_of_point_with_same_non_zero_particle_id / total_num_of_point\n\n2. run DBSCAN. Use different helix unrolling parameters. Adjust EPS so that purity is high (preferably 1). It is ok to break long tracks into smaller ones\n\n3. for each DBSCAN results, for each point construct:\n      feature = length_of_cluster, x,y,x, .... etc ... (other example feature can be x,yz, of closest point, layer_id, direction vector to next/prev layer point ...)\n    \n4. Now for each point, you have:\n        f =  [ feature1 from DBSCAN1, feature2 from DBSCAN2, .... featureN from DBSCANN ]\n\n       the concatenation of N DBSCAN results capture the neighbor characteristics of a point.\n\n5. Train a classifier to map f to y= [ 1, 0,0, ....1, 0 ] (e.g. deep learning or random trees?)\n    dim of y is N. \"1\" at n-th indicate DBSCAN-N results is pure. Use sigmoid because several DBSCAN can be correct\n\n6. Now you know which DBSCAN results is correct. you can construct adjacency grpah or use other methods to fuse small tracks into large ones.\n\n\n\n \n\n",
    "367475": "Dear @Heng Thanks for this discussion! And for links. I am the first time kaggler, learning how to ensemble. Yes, this competition is not the best for the start, but after all this time I invested in it I have to ensemble those models !!! You and @CPMP are here for long, maybe you know if there were discussion and kernels in the previous clustering competitions on the models ensembling for clustering like that? I have a few models, all relatively low scores, and want to ensemble them  to 1. learn how to ensemble and\n 2. to improve my miserable score :)\n\nI found weighted blend by @Sergey Zlobin, and few more https://www.kaggle.com/sergeyzlobin/wanna-blend and here \nand some more here https://www.kaggle.com/vincento/ensemble-your-models-for-better-score\nI found this one useful\nhttps://www.kaggle.com/derrickchua29/ensemble-models-comparison-techniques\n\nI am a first time kaggler with no background in ML/programming, so forgive me if I am wrong. Non of them is really appropriate for this competition and here we need to choose the longest track or so. We can use group by and compare the tracks in different models, and then choose the best one... but how do we make sure we do not assign one hit to two different tracks? How to make sure all hits are also assigned to some tracks? create track 0 where all undefined stuff goes?  I started this competition with googling \"how to install python\" and \"how to use github\",  totally new... sorry if this question is inappropriate, but maybe there is a good material for beginners that i missed to make through in this ensembling for this unusual clustering thing... I believe people with silver-gold medal results know how to ensemble, so only novice people like me struggle with this problem. So, if you could give us few clarifications it would not put you at risk :)  Yes, i saw those python libruary, but it's plug and play, and I need to learn how to do this to be able to do ensembling in future",
    "343864": "Ensembling 5 models gave us a 0.05 boost but then we hit the wall. It seems no matter what the 6th model is, it won't give us more benefit. From our experience, ensembling models with high quality tracks gives us more accuracy than ensembling strong models + weak models. We tried to ensemble models with a cone slicing model, but it degrades after the 5th model, so we had to remove it. Our ensembling code still needs to be improved since @Grzegorz proved it gave him a higher bump.\n\n",
    "340120": "I use an ensembler integrated with the models. Ensembling is partially performed on the fly (while calculations are not finished yet) changing a little bit the state of each model calculations giving them a chance for additional improvement. The boost is ~0.08 for 7 models.\n",
    "339144": "I use largest track to merge, and it seems this is also used in some of the public kernels.",
    "340141": "some notes on \"ensemble clustering\"\n\n@CPMP python scoring function is useful for relabeling in majority voting.\n\nhttps://www.kaggle.com/cpmpml/a-faster-python-scoring-function\n\nhttps://cse.buffalo.edu/~jing/cse601/fa12/materials/clustering_ensemble.pdf\n\nsee also:\n\nhttps://datascience.stackexchange.com/questions/27420/how-to-apply-ensemble-clustering-method\n\n",
    "340121": "@Grzegorz Sionkowski\n \nThanks for the information! \n",
    "367909": ""
  }
}