{
  "id": 60330,
  "title": "merging strategy",
  "url": "/competitions/trackml-particle-identification/discussion/60330",
  "author_name": "",
  "post_date": "2018-07-03T06:55:26.696443700Z",
  "votes": 11,
  "comment_count": 11,
  "views": 0,
  "content": "<p>After 0.8 LB, I still got about 0.05 on the merging game. Currently, I have 13 steps on merging and extension.</p>\n\n<p>For example, apply on the same clusters with different strategy:</p>\n\n<pre><code>merge(sort)-by   score     extend\nrandom           0.6636      x\nlength           0.7681      x\nconfidence       0.8115      x\nmultistage       0.8568      x\nmultistage       0.8839      o\n</code></pre>\n\n<p>We can get 0.22 improvement compare to merging by random.</p>\n\n<p>The whole steps:</p>\n\n<pre><code>step  score   length  tracks  outliers  loss     out-score  remark\n 1    0.4019  11.62   2615    63299     0.0028   0.5953     merge best long tracks\n 2    0.7063  11.19   4954    38265     0.0153   0.2784     merge good long tracks\n 3    0.7227  11.51   4954    36642     0.0174   0.2598     extend\n 4    0.7887  11.23   5594    30848     0.0235   0.1878     merge normal long tracks\n 5    0.8024  11.49   5594    29383     0.0263   0.1713     extend\n 6    0.8095  11.44   5683    28692     0.0277   0.1628     merge bad long tracks\n 7    0.8209  11.66   5683    27416     0.0302   0.1489     extend\n 8    0.8745  10.96   6524    22175     0.0505   0.0750     merge short tracks\n 9    0.8762  11.04   6524    21675     0.0523   0.0715     extend\n10    0.8812  10.37   7288    18115     0.0731   0.0457     merge garbage tracks\n11    0.8834   9.43   8597    12644     0.0960   0.0207     resource recycling\n12    0.8837   9.45   8597    12430     0.0967   0.0196     final extend 1\n13    0.8839   9.46   8597    12357     0.0969   0.0192     final extend 2\n*loss means the lost scores in the tracks\n</code></pre>\n\n<p>So, I can tune parameters on every single step, and observe whether the performance improved. A good improvement, for example, is like more score, less loss, more outliers and so on. If one step is relatively inefficient, I separate it by 2.</p>",
  "messages": [
    {
      "id": "351864",
      "postDate": "07/03/2018 06:55:26",
      "content": "<p>After 0.8 LB, I still got about 0.05 on the merging game. Currently, I have 13 steps on merging and extension.</p>\n\n<p>For example, apply on the same clusters with different strategy:</p>\n\n<pre><code>merge(sort)-by   score     extend\nrandom           0.6636      x\nlength           0.7681      x\nconfidence       0.8115      x\nmultistage       0.8568      x\nmultistage       0.8839      o\n</code></pre>\n\n<p>We can get 0.22 improvement compare to merging by random.</p>\n\n<p>The whole steps:</p>\n\n<pre><code>step  score   length  tracks  outliers  loss     out-score  remark\n 1    0.4019  11.62   2615    63299     0.0028   0.5953     merge best long tracks\n 2    0.7063  11.19   4954    38265     0.0153   0.2784     merge good long tracks\n 3    0.7227  11.51   4954    36642     0.0174   0.2598     extend\n 4    0.7887  11.23   5594    30848     0.0235   0.1878     merge normal long tracks\n 5    0.8024  11.49   5594    29383     0.0263   0.1713     extend\n 6    0.8095  11.44   5683    28692     0.0277   0.1628     merge bad long tracks\n 7    0.8209  11.66   5683    27416     0.0302   0.1489     extend\n 8    0.8745  10.96   6524    22175     0.0505   0.0750     merge short tracks\n 9    0.8762  11.04   6524    21675     0.0523   0.0715     extend\n10    0.8812  10.37   7288    18115     0.0731   0.0457     merge garbage tracks\n11    0.8834   9.43   8597    12644     0.0960   0.0207     resource recycling\n12    0.8837   9.45   8597    12430     0.0967   0.0196     final extend 1\n13    0.8839   9.46   8597    12357     0.0969   0.0192     final extend 2\n*loss means the lost scores in the tracks\n</code></pre>\n\n<p>So, I can tune parameters on every single step, and observe whether the performance improved. A good improvement, for example, is like more score, less loss, more outliers and so on. If one step is relatively inefficient, I separate it by 2.</p>",
      "rawMarkdown": "After 0.8 LB, I still got about 0.05 on the merging game. Currently, I have 13 steps on merging and extension.\n\nFor example, apply on the same clusters with different strategy:\n\n    merge(sort)-by   score     extend\n    random           0.6636      x\n    length           0.7681      x\n    confidence       0.8115      x\n    multistage       0.8568      x\n    multistage       0.8839      o\n\nWe can get 0.22 improvement compare to merging by random.\n\nThe whole steps:\n\n    step  score   length  tracks  outliers  loss     out-score  remark\n     1    0.4019  11.62   2615    63299     0.0028   0.5953     merge best long tracks\n     2    0.7063  11.19   4954    38265     0.0153   0.2784     merge good long tracks\n     3    0.7227  11.51   4954    36642     0.0174   0.2598     extend\n     4    0.7887  11.23   5594    30848     0.0235   0.1878     merge normal long tracks\n     5    0.8024  11.49   5594    29383     0.0263   0.1713     extend\n     6    0.8095  11.44   5683    28692     0.0277   0.1628     merge bad long tracks\n     7    0.8209  11.66   5683    27416     0.0302   0.1489     extend\n     8    0.8745  10.96   6524    22175     0.0505   0.0750     merge short tracks\n     9    0.8762  11.04   6524    21675     0.0523   0.0715     extend\n    10    0.8812  10.37   7288    18115     0.0731   0.0457     merge garbage tracks\n    11    0.8834   9.43   8597    12644     0.0960   0.0207     resource recycling\n    12    0.8837   9.45   8597    12430     0.0967   0.0196     final extend 1\n    13    0.8839   9.46   8597    12357     0.0969   0.0192     final extend 2\n    *loss means the lost scores in the tracks\n\nSo, I can tune parameters on every single step, and observe whether the performance improved. A good improvement, for example, is like more score, less loss, more outliers and so on. If one step is relatively inefficient, I separate it by 2.",
      "votes": null
    },
    {
      "id": "351877",
      "postDate": "07/03/2018 07:49:08",
      "content": "<p><a href=\"/outrunner\">@outrunner</a>, Thank you very much for sharing your approach.\nI do not know what kind of merging do you use, but after merging/ensembling the best and good tracks (step 1 and 2) you may obtain some short tracks. As I understand you prefer to extend them in step 3. My approach is different, but by analogy I would separate them after step 2 and process together with the other short tracks (step 8). Maybe this step (separation of short tracks) is so obvious that you did not mentioned it.</p>",
      "rawMarkdown": "outrunner, Thank you very much for sharing your approach.\nI do not know what kind of merging do you use, but after merging/ensembling the best and good tracks (step 1 and 2) you may obtain some short tracks. As I understand you prefer to extend them in step 3. My approach is different, but by analogy I would separate them after step 2 and process together with the other short tracks (step 8). Maybe this step (separation of short tracks) is so obvious that you did not mentioned it.",
      "votes": null
    },
    {
      "id": "351890",
      "postDate": "07/03/2018 08:28:26",
      "content": "<p>Thanks @Grzegorz Sionkowski, we do the same if I'm not misunderstanding. Step 3 only extend the tracks assigned on previous two steps. The rest (good) short tracks will remain to the step 8.</p>",
      "rawMarkdown": "Thanks @Grzegorz Sionkowski, we do the same if I'm not misunderstanding. Step 3 only extend the tracks assigned on previous two steps. The rest (good) short tracks will remain to the step 8.",
      "votes": null
    },
    {
      "id": "351898",
      "postDate": "07/03/2018 08:45:55",
      "content": "<p><a href=\"/outrunner\">@outrunner</a>, To be finally well understood. Talking about short tracks I mean the effect of step 1 and 2,  i.e. the short tracks being the effect of ensembling best and good long tracks, not the good short tracks obtained before the step 1. If you use merging acting as connecting only, you will not obtain the tracks I mean.</p>",
      "rawMarkdown": "outrunner, To be finally well understood. Talking about short tracks I mean the effect of step 1 and 2,  i.e. the short tracks being the effect of ensembling best and good long tracks, not the good short tracks obtained before the step 1. If you use merging acting as connecting only, you will not obtain the tracks I mean.",
      "votes": null
    },
    {
      "id": "351910",
      "postDate": "07/03/2018 09:13:50",
      "content": "<p>ok, hope I do understand this time. The effected track (long but some hits are occupied by previous merging tracks) will be short track too.</p>",
      "rawMarkdown": "ok, hope I do understand this time. The effected track (long but some hits are occupied by previous merging tracks) will be short track too.",
      "votes": null
    },
    {
      "id": "351978",
      "postDate": "07/03/2018 12:51:37",
      "content": "<p><a href=\"/outrunner\">@outrunner</a>, If the sequence of merging the tracks plays so important role, I guess your merging method is more similar to the simple method used in the public kernels (a track cannot  become longer) than to any sophisticated one.</p>",
      "rawMarkdown": "outrunner, If the sequence of merging the tracks plays so important role, I guess your merging method is more similar to the simple method used in the public kernels (a track cannot  become longer) than to any sophisticated one.",
      "votes": null
    },
    {
      "id": "352024",
      "postDate": "07/03/2018 14:00:29",
      "content": "<p>I can not find the kernel, but I think you are right. My code:</p>\n\n<pre><code>for id in merge_ids:\n    track = tracks_all[id]\n    track = track[np.where(tracks[track]==0)[0]]\n\n    if len(track) &gt;= min_length:\n        track_id = track_id + 1  \n        tracks[track] = track_id\n    else:\n        short_tracks.append(id)\n</code></pre>",
      "rawMarkdown": "I can not find the kernel, but I think you are right. My code:\n\n    for id in merge_ids:\n        track = tracks_all[id]\n        track = track[np.where(tracks[track]==0)[0]]\n\n        if len(track) &gt;= min_length:\n            track_id = track_id + 1  \n            tracks[track] = track_id\n        else:\n            short_tracks.append(id)",
      "votes": null
    },
    {
      "id": "352831",
      "postDate": "07/05/2018 08:11:54",
      "content": "<p><a href=\"/outrunner\">@outrunner</a>, Could you tell us how do you evaluate the quality of tracks (best, good, normal, bad)? Internally, on the base of properties of hits in the track, or \"externally\", on the base of quality of tracks of other event obtained the same way?</p>",
      "rawMarkdown": "outrunner, Could you tell us how do you evaluate the quality of tracks (best, good, normal, bad)? Internally, on the base of properties of hits in the track, or \"externally\", on the base of quality of tracks of other event obtained the same way?",
      "votes": null
    },
    {
      "id": "352874",
      "postDate": "07/05/2018 10:22:12",
      "content": "<p><a href=\"/outrunner\">@outrunner</a>, Some of us (at least one) use more sophisticated merging method, but majority of us cannot obtain the score 0.7+ merging the tracks sorted by length only. You have excellent track candidates.</p>",
      "rawMarkdown": "outrunner, Some of us (at least one) use more sophisticated merging method, but majority of us cannot obtain the score 0.7+ merging the tracks sorted by length only. You have excellent track candidates.",
      "votes": null
    },
    {
      "id": "352911",
      "postDate": "07/05/2018 12:14:44",
      "content": "<p>The confidence score is based on statistics, event-wise. As you mentioned, the most important thing is the quality of track candidates.</p>",
      "rawMarkdown": "The confidence score is based on statistics, event-wise. As you mentioned, the most important thing is the quality of track candidates.",
      "votes": null
    },
    {
      "id": "354058",
      "postDate": "07/08/2018 15:49:20",
      "content": "<p><a href=\"/outrunner\">@outrunner</a>, you can still fight for about 2-7% lost in out-score by attaching the outliers to the tracks of even number of hits. <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60638#354053\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60638#354053</a></p>",
      "rawMarkdown": "outrunner, you can still fight for about 2-7% lost in out-score by attaching the outliers to the tracks of even number of hits. https://www.kaggle.com/c/trackml-particle-identification/discussion/60638#354053",
      "votes": null
    },
    {
      "id": "354063",
      "postDate": "07/08/2018 16:19:54",
      "content": "<p>@Grzegorz Sionkowski, this is the most useful information, thanks.</p>",
      "rawMarkdown": "Grzegorz Sionkowski, this is the most useful information, thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 351877,
      "author_name": "sionek",
      "author_url": "",
      "post_date": "07/03/2018 07:49:08",
      "content": "<p><a href=\"/outrunner\">@outrunner</a>, Thank you very much for sharing your approach.\nI do not know what kind of merging do you use, but after merging/ensembling the best and good tracks (step 1 and 2) you may obtain some short tracks. As I understand you prefer to extend them in step 3. My approach is different, but by analogy I would separate them after step 2 and process together with the other short tracks (step 8). Maybe this step (separation of short tracks) is so obvious that you did not mentioned it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 351890,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "07/03/2018 08:28:26",
          "content": "<p>Thanks @Grzegorz Sionkowski, we do the same if I'm not misunderstanding. Step 3 only extend the tracks assigned on previous two steps. The rest (good) short tracks will remain to the step 8.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 351898,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "07/03/2018 08:45:55",
          "content": "<p><a href=\"/outrunner\">@outrunner</a>, To be finally well understood. Talking about short tracks I mean the effect of step 1 and 2,  i.e. the short tracks being the effect of ensembling best and good long tracks, not the good short tracks obtained before the step 1. If you use merging acting as connecting only, you will not obtain the tracks I mean.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 351910,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "07/03/2018 09:13:50",
          "content": "<p>ok, hope I do understand this time. The effected track (long but some hits are occupied by previous merging tracks) will be short track too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 351978,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "07/03/2018 12:51:37",
          "content": "<p><a href=\"/outrunner\">@outrunner</a>, If the sequence of merging the tracks plays so important role, I guess your merging method is more similar to the simple method used in the public kernels (a track cannot  become longer) than to any sophisticated one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 352024,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "07/03/2018 14:00:29",
          "content": "<p>I can not find the kernel, but I think you are right. My code:</p>\n\n<pre><code>for id in merge_ids:\n    track = tracks_all[id]\n    track = track[np.where(tracks[track]==0)[0]]\n\n    if len(track) &gt;= min_length:\n        track_id = track_id + 1  \n        tracks[track] = track_id\n    else:\n        short_tracks.append(id)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 352874,
          "author_name": "sionek",
          "author_url": "",
          "post_date": "07/05/2018 10:22:12",
          "content": "<p><a href=\"/outrunner\">@outrunner</a>, Some of us (at least one) use more sophisticated merging method, but majority of us cannot obtain the score 0.7+ merging the tracks sorted by length only. You have excellent track candidates.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 352831,
      "author_name": "sionek",
      "author_url": "",
      "post_date": "07/05/2018 08:11:54",
      "content": "<p><a href=\"/outrunner\">@outrunner</a>, Could you tell us how do you evaluate the quality of tracks (best, good, normal, bad)? Internally, on the base of properties of hits in the track, or \"externally\", on the base of quality of tracks of other event obtained the same way?</p>",
      "votes": null,
      "replies": [
        {
          "id": 352911,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "07/05/2018 12:14:44",
          "content": "<p>The confidence score is based on statistics, event-wise. As you mentioned, the most important thing is the quality of track candidates.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 354058,
      "author_name": "sionek",
      "author_url": "",
      "post_date": "07/08/2018 15:49:20",
      "content": "<p><a href=\"/outrunner\">@outrunner</a>, you can still fight for about 2-7% lost in out-score by attaching the outliers to the tracks of even number of hits. <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/60638#354053\">https://www.kaggle.com/c/trackml-particle-identification/discussion/60638#354053</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 354063,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "07/08/2018 16:19:54",
          "content": "<p>@Grzegorz Sionkowski, this is the most useful information, thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "351864": "After 0.8 LB, I still got about 0.05 on the merging game. Currently, I have 13 steps on merging and extension.\n\nFor example, apply on the same clusters with different strategy:\n\n    merge(sort)-by   score     extend\n    random           0.6636      x\n    length           0.7681      x\n    confidence       0.8115      x\n    multistage       0.8568      x\n    multistage       0.8839      o\n\nWe can get 0.22 improvement compare to merging by random.\n\nThe whole steps:\n\n    step  score   length  tracks  outliers  loss     out-score  remark\n     1    0.4019  11.62   2615    63299     0.0028   0.5953     merge best long tracks\n     2    0.7063  11.19   4954    38265     0.0153   0.2784     merge good long tracks\n     3    0.7227  11.51   4954    36642     0.0174   0.2598     extend\n     4    0.7887  11.23   5594    30848     0.0235   0.1878     merge normal long tracks\n     5    0.8024  11.49   5594    29383     0.0263   0.1713     extend\n     6    0.8095  11.44   5683    28692     0.0277   0.1628     merge bad long tracks\n     7    0.8209  11.66   5683    27416     0.0302   0.1489     extend\n     8    0.8745  10.96   6524    22175     0.0505   0.0750     merge short tracks\n     9    0.8762  11.04   6524    21675     0.0523   0.0715     extend\n    10    0.8812  10.37   7288    18115     0.0731   0.0457     merge garbage tracks\n    11    0.8834   9.43   8597    12644     0.0960   0.0207     resource recycling\n    12    0.8837   9.45   8597    12430     0.0967   0.0196     final extend 1\n    13    0.8839   9.46   8597    12357     0.0969   0.0192     final extend 2\n    *loss means the lost scores in the tracks\n\nSo, I can tune parameters on every single step, and observe whether the performance improved. A good improvement, for example, is like more score, less loss, more outliers and so on. If one step is relatively inefficient, I separate it by 2.",
    "351877": "outrunner, Thank you very much for sharing your approach.\nI do not know what kind of merging do you use, but after merging/ensembling the best and good tracks (step 1 and 2) you may obtain some short tracks. As I understand you prefer to extend them in step 3. My approach is different, but by analogy I would separate them after step 2 and process together with the other short tracks (step 8). Maybe this step (separation of short tracks) is so obvious that you did not mentioned it.",
    "351890": "Thanks @Grzegorz Sionkowski, we do the same if I'm not misunderstanding. Step 3 only extend the tracks assigned on previous two steps. The rest (good) short tracks will remain to the step 8.",
    "351898": "outrunner, To be finally well understood. Talking about short tracks I mean the effect of step 1 and 2,  i.e. the short tracks being the effect of ensembling best and good long tracks, not the good short tracks obtained before the step 1. If you use merging acting as connecting only, you will not obtain the tracks I mean.",
    "351910": "ok, hope I do understand this time. The effected track (long but some hits are occupied by previous merging tracks) will be short track too.",
    "351978": "outrunner, If the sequence of merging the tracks plays so important role, I guess your merging method is more similar to the simple method used in the public kernels (a track cannot  become longer) than to any sophisticated one.",
    "352024": "I can not find the kernel, but I think you are right. My code:\n\n    for id in merge_ids:\n        track = tracks_all[id]\n        track = track[np.where(tracks[track]==0)[0]]\n\n        if len(track) &gt;= min_length:\n            track_id = track_id + 1  \n            tracks[track] = track_id\n        else:\n            short_tracks.append(id)",
    "352831": "outrunner, Could you tell us how do you evaluate the quality of tracks (best, good, normal, bad)? Internally, on the base of properties of hits in the track, or \"externally\", on the base of quality of tracks of other event obtained the same way?",
    "352874": "outrunner, Some of us (at least one) use more sophisticated merging method, but majority of us cannot obtain the score 0.7+ merging the tracks sorted by length only. You have excellent track candidates.",
    "352911": "The confidence score is based on statistics, event-wise. As you mentioned, the most important thing is the quality of track candidates.",
    "354058": "outrunner, you can still fight for about 2-7% lost in out-score by attaching the outliers to the tracks of even number of hits. https://www.kaggle.com/c/trackml-particle-identification/discussion/60638#354053",
    "354063": "Grzegorz Sionkowski, this is the most useful information, thanks."
  },
  "source": "meta"
}