{
  "id": 63303,
  "title": "11th place solution",
  "url": "/competitions/trackml-particle-identification/writeups/andrea-lonza-11th-place-solution",
  "author_name": "",
  "post_date": "2018-08-26T14:01:59.257Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi,\nit's been a very interesting competition and as always these challenges are a very good place to learn. \nI tried many different approaches but eventually, as many of you, I used clustering (dbscan) on unrolled helix with z shifting and track extension.</p>\n\n<p>I don't want to explain here the tricks that I used on the clustering part (some of them has already been shared by @yuval and @CPMP) but I do want to explain the feature that allowed me to increase the score from approx. 0.68 (obtained using the clustering algo) to about 0.76, namely a supervised track extension. </p>\n\n<p><strong>Supervised track extension</strong>\nThe base code is similar to the one shared by @HengCherKeng where I integrated a gradient boosting tree (LightGBM) to establish if a given hit belongs to a given track. As input it takes approx. 60 features of the track (constructed using the clustering algo) and the proposed hit, and it outputs the probability of the hit to belong to the track. In this way, I could take into consideration a larger number of potential hits and let the algorithm decide which one is the best candidate.\nI trained the decision tree over only 10 events and I used pretty naive features, meaning that it can be improved much more.</p>\n\n<p>From a performance point of view, the clustering algorithm takes about 1h per event on a single core instead, the extension algorithm takes about 20min (at inference time). </p>\n\n<p>If you are interested, in the next days I will share the code.</p>\n\n<p>Thanks to all of you who shared your ideas during the competition, cheers!</p>\n\n<p>NEWS!</p>\n\n<p>The code and documentation are available on <a href=\"https://github.com/andri27-ts/GoldTrackML\">https://github.com/andri27-ts/GoldTrackML</a> </p>",
  "messages": [
    {
      "id": "370298",
      "postDate": "08/14/2018 15:30:07",
      "content": "<p>Hi,\nit's been a very interesting competition and as always these challenges are a very good place to learn. \nI tried many different approaches but eventually, as many of you, I used clustering (dbscan) on unrolled helix with z shifting and track extension.</p>\n\n<p>I don't want to explain here the tricks that I used on the clustering part (some of them has already been shared by @yuval and @CPMP) but I do want to explain the feature that allowed me to increase the score from approx. 0.68 (obtained using the clustering algo) to about 0.76, namely a supervised track extension. </p>\n\n<p><strong>Supervised track extension</strong>\nThe base code is similar to the one shared by @HengCherKeng where I integrated a gradient boosting tree (LightGBM) to establish if a given hit belongs to a given track. As input it takes approx. 60 features of the track (constructed using the clustering algo) and the proposed hit, and it outputs the probability of the hit to belong to the track. In this way, I could take into consideration a larger number of potential hits and let the algorithm decide which one is the best candidate.\nI trained the decision tree over only 10 events and I used pretty naive features, meaning that it can be improved much more.</p>\n\n<p>From a performance point of view, the clustering algorithm takes about 1h per event on a single core instead, the extension algorithm takes about 20min (at inference time). </p>\n\n<p>If you are interested, in the next days I will share the code.</p>\n\n<p>Thanks to all of you who shared your ideas during the competition, cheers!</p>\n\n<p>NEWS!</p>\n\n<p>The code and documentation are available on <a href=\"https://github.com/andri27-ts/GoldTrackML\">https://github.com/andri27-ts/GoldTrackML</a> </p>",
      "rawMarkdown": "Hi,\nit's been a very interesting competition and as always these challenges are a very good place to learn. \nI tried many different approaches but eventually, as many of you, I used clustering (dbscan) on unrolled helix with z shifting and track extension.\n\nI don't want to explain here the tricks that I used on the clustering part (some of them has already been shared by @yuval and @CPMP) but I do want to explain the feature that allowed me to increase the score from approx. 0.68 (obtained using the clustering algo) to about 0.76, namely a supervised track extension. \n\n**Supervised track extension**\nThe base code is similar to the one shared by @HengCherKeng where I integrated a gradient boosting tree (LightGBM) to establish if a given hit belongs to a given track. As input it takes approx. 60 features of the track (constructed using the clustering algo) and the proposed hit, and it outputs the probability of the hit to belong to the track. In this way, I could take into consideration a larger number of potential hits and let the algorithm decide which one is the best candidate.\nI trained the decision tree over only 10 events and I used pretty naive features, meaning that it can be improved much more.\n\nFrom a performance point of view, the clustering algorithm takes about 1h per event on a single core instead, the extension algorithm takes about 20min (at inference time). \n\nIf you are interested, in the next days I will share the code.\n\nThanks to all of you who shared your ideas during the competition, cheers!\n\nNEWS!\n\nThe code and documentation are available on https://github.com/andri27-ts/GoldTrackML",
      "votes": null
    },
    {
      "id": "370306",
      "postDate": "08/14/2018 15:43:01",
      "content": "<p>Thanks for sharing, and congrats on the result.  </p>",
      "rawMarkdown": "Thanks for sharing, and congrats on the result.",
      "votes": null
    },
    {
      "id": "370631",
      "postDate": "08/15/2018 07:20:08",
      "content": "<p>@Andrea, the supervised track extension idea is pretty neat. I'm very interested in a supervised solution. I didn't think of using it for track extension, but seeding and track fitting, but track fitting and track extension are actually the same thing. It'd be great if you share your code. :) Great job!! We always had trouble catching up with your score even at the beginning of this competition. Of course at the end of this competition too. ;)</p>",
      "rawMarkdown": "Andrea, the supervised track extension idea is pretty neat. I'm very interested in a supervised solution. I didn't think of using it for track extension, but seeding and track fitting, but track fitting and track extension are actually the same thing. It'd be great if you share your code. :) Great job!! We always had trouble catching up with your score even at the beginning of this competition. Of course at the end of this competition too. ;)",
      "votes": null
    },
    {
      "id": "371029",
      "postDate": "08/15/2018 21:04:05",
      "content": "<p>I've tried boosting trees instead of Heng's code. But probably I had a few features.</p>",
      "rawMarkdown": "I've tried boosting trees instead of Heng's code. But probably I had a few features.",
      "votes": null
    },
    {
      "id": "371233",
      "postDate": "08/16/2018 09:57:50",
      "content": "<p>Yes, I agree that track fitting and track extension are basically the same concepts, but I think that can be hard to use supervision methods for track seeding. In this case you have to redefine the problem because it's not more a classification task as for track extension. Thanks, and congratulation to you too.. and wasn't my idea to catch you the last day but I didn't have time the previous weeks</p>",
      "rawMarkdown": "Yes, I agree that track fitting and track extension are basically the same concepts, but I think that can be hard to use supervision methods for track seeding. In this case you have to redefine the problem because it's not more a classification task as for track extension. Thanks, and congratulation to you too.. and wasn't my idea to catch you the last day but I didn't have time the previous weeks",
      "votes": null
    },
    {
      "id": "371235",
      "postDate": "08/16/2018 10:03:30",
      "content": "<p>Technically at the end, I used gradient boosting tree (LightGBM). If you didn't use Heng's code, how did you choose the potential hits that could belong to a given track?</p>",
      "rawMarkdown": "Technically at the end, I used gradient boosting tree (LightGBM). If you didn't use Heng's code, how did you choose the potential hits that could belong to a given track?",
      "votes": null
    },
    {
      "id": "371246",
      "postDate": "08/16/2018 10:38:35",
      "content": "<p>I tried to predict a next hit by 3 last hits of a track. Then we can choose a closest hit to the predicted.</p>",
      "rawMarkdown": "I tried to predict a next hit by 3 last hits of a track. Then we can choose a closest hit to the predicted.",
      "votes": null
    },
    {
      "id": "371255",
      "postDate": "08/16/2018 10:51:48",
      "content": "<p>@Andrea, for track seeding, <a href=\"/outrunner\">@outrunner</a> made it, oh wait, I think his NN model finds more than just seeds :)  </p>\n\n<p>You may not believe it, we're actually very happy that such an innovative supervised learning approach beat our score at the end after having read your post. Every new idea is a contribution to CERN.</p>",
      "rawMarkdown": "Andrea, for track seeding, @outrunner made it, oh wait, I think his NN model finds more than just seeds :)  \n\nYou may not believe it, we're actually very happy that such an innovative supervised learning approach beat our score at the end after having read your post. Every new idea is a contribution to CERN.",
      "votes": null
    },
    {
      "id": "371277",
      "postDate": "08/16/2018 12:37:00",
      "content": "<p>Yes, sure. I was talking more about the difficulty from a computational time point of view.</p>",
      "rawMarkdown": "Yes, sure. I was talking more about the difficulty from a computational time point of view.",
      "votes": null
    },
    {
      "id": "375035",
      "postDate": "08/24/2018 12:35:25",
      "content": "<p>The code and the documentation are on Github <a href=\"https://github.com/andri27-ts/GoldTrackML\">https://github.com/andri27-ts/GoldTrackML</a> !!\nHave a nice day!</p>",
      "rawMarkdown": "The code and the documentation are on Github https://github.com/andri27-ts/GoldTrackML !!\nHave a nice day!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 370306,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/14/2018 15:43:01",
      "content": "<p>Thanks for sharing, and congrats on the result.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 370631,
      "author_name": "nicolefinnie",
      "author_url": "",
      "post_date": "08/15/2018 07:20:08",
      "content": "<p>@Andrea, the supervised track extension idea is pretty neat. I'm very interested in a supervised solution. I didn't think of using it for track extension, but seeding and track fitting, but track fitting and track extension are actually the same thing. It'd be great if you share your code. :) Great job!! We always had trouble catching up with your score even at the beginning of this competition. Of course at the end of this competition too. ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 371233,
          "author_name": "andri27",
          "author_url": "",
          "post_date": "08/16/2018 09:57:50",
          "content": "<p>Yes, I agree that track fitting and track extension are basically the same concepts, but I think that can be hard to use supervision methods for track seeding. In this case you have to redefine the problem because it's not more a classification task as for track extension. Thanks, and congratulation to you too.. and wasn't my idea to catch you the last day but I didn't have time the previous weeks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371255,
          "author_name": "nicolefinnie",
          "author_url": "",
          "post_date": "08/16/2018 10:51:48",
          "content": "<p>@Andrea, for track seeding, <a href=\"/outrunner\">@outrunner</a> made it, oh wait, I think his NN model finds more than just seeds :)  </p>\n\n<p>You may not believe it, we're actually very happy that such an innovative supervised learning approach beat our score at the end after having read your post. Every new idea is a contribution to CERN.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371277,
          "author_name": "andri27",
          "author_url": "",
          "post_date": "08/16/2018 12:37:00",
          "content": "<p>Yes, sure. I was talking more about the difficulty from a computational time point of view.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 371029,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "08/15/2018 21:04:05",
      "content": "<p>I've tried boosting trees instead of Heng's code. But probably I had a few features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 371235,
          "author_name": "andri27",
          "author_url": "",
          "post_date": "08/16/2018 10:03:30",
          "content": "<p>Technically at the end, I used gradient boosting tree (LightGBM). If you didn't use Heng's code, how did you choose the potential hits that could belong to a given track?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 371246,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/16/2018 10:38:35",
          "content": "<p>I tried to predict a next hit by 3 last hits of a track. Then we can choose a closest hit to the predicted.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 375035,
      "author_name": "andri27",
      "author_url": "",
      "post_date": "08/24/2018 12:35:25",
      "content": "<p>The code and the documentation are on Github <a href=\"https://github.com/andri27-ts/GoldTrackML\">https://github.com/andri27-ts/GoldTrackML</a> !!\nHave a nice day!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "370298": "Hi,\nit's been a very interesting competition and as always these challenges are a very good place to learn. \nI tried many different approaches but eventually, as many of you, I used clustering (dbscan) on unrolled helix with z shifting and track extension.\n\nI don't want to explain here the tricks that I used on the clustering part (some of them has already been shared by @yuval and @CPMP) but I do want to explain the feature that allowed me to increase the score from approx. 0.68 (obtained using the clustering algo) to about 0.76, namely a supervised track extension. \n\n**Supervised track extension**\nThe base code is similar to the one shared by @HengCherKeng where I integrated a gradient boosting tree (LightGBM) to establish if a given hit belongs to a given track. As input it takes approx. 60 features of the track (constructed using the clustering algo) and the proposed hit, and it outputs the probability of the hit to belong to the track. In this way, I could take into consideration a larger number of potential hits and let the algorithm decide which one is the best candidate.\nI trained the decision tree over only 10 events and I used pretty naive features, meaning that it can be improved much more.\n\nFrom a performance point of view, the clustering algorithm takes about 1h per event on a single core instead, the extension algorithm takes about 20min (at inference time). \n\nIf you are interested, in the next days I will share the code.\n\nThanks to all of you who shared your ideas during the competition, cheers!\n\nNEWS!\n\nThe code and documentation are available on https://github.com/andri27-ts/GoldTrackML",
    "370306": "Thanks for sharing, and congrats on the result.",
    "370631": "Andrea, the supervised track extension idea is pretty neat. I'm very interested in a supervised solution. I didn't think of using it for track extension, but seeding and track fitting, but track fitting and track extension are actually the same thing. It'd be great if you share your code. :) Great job!! We always had trouble catching up with your score even at the beginning of this competition. Of course at the end of this competition too. ;)",
    "371029": "I've tried boosting trees instead of Heng's code. But probably I had a few features.",
    "371233": "Yes, I agree that track fitting and track extension are basically the same concepts, but I think that can be hard to use supervision methods for track seeding. In this case you have to redefine the problem because it's not more a classification task as for track extension. Thanks, and congratulation to you too.. and wasn't my idea to catch you the last day but I didn't have time the previous weeks",
    "371235": "Technically at the end, I used gradient boosting tree (LightGBM). If you didn't use Heng's code, how did you choose the potential hits that could belong to a given track?",
    "371246": "I tried to predict a next hit by 3 last hits of a track. Then we can choose a closest hit to the predicted.",
    "371255": "Andrea, for track seeding, @outrunner made it, oh wait, I think his NN model finds more than just seeds :)  \n\nYou may not believe it, we're actually very happy that such an innovative supervised learning approach beat our score at the end after having read your post. Every new idea is a contribution to CERN.",
    "371277": "Yes, sure. I was talking more about the difficulty from a computational time point of view.",
    "375035": "The code and the documentation are on Github https://github.com/andri27-ts/GoldTrackML !!\nHave a nice day!"
  },
  "source": "meta"
}