{
  "id": 56268,
  "title": "Solution to Duplicate Problem by Reverse Engineering (0.0005 Boost)",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/k-a-c-solution-to-duplicate-problem-by-reverse-eng",
  "author_name": "",
  "post_date": "2018-05-08T06:26:03.709757Z",
  "votes": 81,
  "comment_count": 19,
  "views": 0,
  "content": "<p>As you have all noticed, there were duplicate samples with different labels. I will chronologically explain how I found a solution for it. </p>\n\n<p>In the middle of competition, I think I have found time_to_next_click on [ip, app, device, os] feature earlier than public kernels and most of the LB. There was a discussion about the fact that click ids are not sorted. @Bojan stated that click ids are leaky in some TalkingData competitions. At that time, I didn't take this comment serious, and I even made a joke about IP%27. Maybe, some people had believed due to my high score which is actually thanks to relatively early discovery of time_to_next_click feature but not leak.</p>\n\n<p>Then I have put serious thinking effort on two things: Why time_to_next_click works best in [ip, app, device, os] group but nothing else like channel? Why there are duplicates with different labels?</p>\n\n<p>My theory is TalkingData has 2 tables: Clicks table C(click_id, ip, app, device, os, channel, time) and Downloads table D(ip, app, device, os, time). So, the labels we see are not necessarily real downloads. This is how TalkingData maps D to C. If you check download times, some of the apps are downloaded 12 hours after they are clicked. This is not realistic. So a download in D is mapped to <strong>the latest</strong> click on C which has the same <strong>(ip, app, device, os)</strong> . <strong>channel</strong> is not important because there is no such thing as channel of a download. Then I noticed that we don't know which click is the latest if they are in the same second. I thought I can try ordering them by click_id, but this information is not available in training set. Then I thought if my logic is correct I can order the training set by (time, is_attributed) if and only if I have one time delta feature which is on (ip, app, device, os).</p>\n\n<p>So, sorting the training data by <strong>(time, is_attributed)</strong>, and the test data by <strong>(time, click_id)</strong> gives +0.0005 boost on public LB if the time delta feature is used. It doesn't work with other time deltas because sorting by (time, is_attributed) is not valid for them and leak wrong information.</p>\n\n<p>Maybe others have better solution for this, but for me it was a big discovery. Since it was not a random discovery but serious thinking effort, I couldn't feel comfortable to share this during the competition.</p>",
  "messages": [
    {
      "id": "325119",
      "postDate": "05/08/2018 06:26:03",
      "content": "<p>As you have all noticed, there were duplicate samples with different labels. I will chronologically explain how I found a solution for it. </p>\n\n<p>In the middle of competition, I think I have found time_to_next_click on [ip, app, device, os] feature earlier than public kernels and most of the LB. There was a discussion about the fact that click ids are not sorted. @Bojan stated that click ids are leaky in some TalkingData competitions. At that time, I didn't take this comment serious, and I even made a joke about IP%27. Maybe, some people had believed due to my high score which is actually thanks to relatively early discovery of time_to_next_click feature but not leak.</p>\n\n<p>Then I have put serious thinking effort on two things: Why time_to_next_click works best in [ip, app, device, os] group but nothing else like channel? Why there are duplicates with different labels?</p>\n\n<p>My theory is TalkingData has 2 tables: Clicks table C(click_id, ip, app, device, os, channel, time) and Downloads table D(ip, app, device, os, time). So, the labels we see are not necessarily real downloads. This is how TalkingData maps D to C. If you check download times, some of the apps are downloaded 12 hours after they are clicked. This is not realistic. So a download in D is mapped to <strong>the latest</strong> click on C which has the same <strong>(ip, app, device, os)</strong> . <strong>channel</strong> is not important because there is no such thing as channel of a download. Then I noticed that we don't know which click is the latest if they are in the same second. I thought I can try ordering them by click_id, but this information is not available in training set. Then I thought if my logic is correct I can order the training set by (time, is_attributed) if and only if I have one time delta feature which is on (ip, app, device, os).</p>\n\n<p>So, sorting the training data by <strong>(time, is_attributed)</strong>, and the test data by <strong>(time, click_id)</strong> gives +0.0005 boost on public LB if the time delta feature is used. It doesn't work with other time deltas because sorting by (time, is_attributed) is not valid for them and leak wrong information.</p>\n\n<p>Maybe others have better solution for this, but for me it was a big discovery. Since it was not a random discovery but serious thinking effort, I couldn't feel comfortable to share this during the competition.</p>",
      "rawMarkdown": "As you have all noticed, there were duplicate samples with different labels. I will chronologically explain how I found a solution for it. \n\nIn the middle of competition, I think I have found time_to_next_click on [ip, app, device, os] feature earlier than public kernels and most of the LB. There was a discussion about the fact that click ids are not sorted. @Bojan stated that click ids are leaky in some TalkingData competitions. At that time, I didn't take this comment serious, and I even made a joke about IP%27. Maybe, some people had believed due to my high score which is actually thanks to relatively early discovery of time_to_next_click feature but not leak.\n\nThen I have put serious thinking effort on two things: Why time_to_next_click works best in [ip, app, device, os] group but nothing else like channel? Why there are duplicates with different labels?\n\nMy theory is TalkingData has 2 tables: Clicks table C(click_id, ip, app, device, os, channel, time) and Downloads table D(ip, app, device, os, time). So, the labels we see are not necessarily real downloads. This is how TalkingData maps D to C. If you check download times, some of the apps are downloaded 12 hours after they are clicked. This is not realistic. So a download in D is mapped to **the latest** click on C which has the same **(ip, app, device, os)** . **channel** is not important because there is no such thing as channel of a download. Then I noticed that we don't know which click is the latest if they are in the same second. I thought I can try ordering them by click_id, but this information is not available in training set. Then I thought if my logic is correct I can order the training set by (time, is_attributed) if and only if I have one time delta feature which is on (ip, app, device, os).\n\nSo, sorting the training data by **(time, is_attributed)**, and the test data by **(time, click_id)** gives +0.0005 boost on public LB if the time delta feature is used. It doesn't work with other time deltas because sorting by (time, is_attributed) is not valid for them and leak wrong information.\n\nMaybe others have better solution for this, but for me it was a big discovery. Since it was not a random discovery but serious thinking effort, I couldn't feel comfortable to share this during the competition.",
      "votes": null
    },
    {
      "id": "325142",
      "postDate": "05/08/2018 06:59:31",
      "content": "<p>Great analysis,\nThanks for sharing.</p>",
      "rawMarkdown": "Great analysis,\nThanks for sharing.",
      "votes": null
    },
    {
      "id": "325152",
      "postDate": "05/08/2018 07:15:43",
      "content": "<p>Very impressive discovery! You guys are soooooooooo smart. Thanks for sharing. BTW, does it mean that <strong>channel</strong> is kind of misleading?</p>",
      "rawMarkdown": "Very impressive discovery! You guys are soooooooooo smart. Thanks for sharing. BTW, does it mean that **channel** is kind of misleading?",
      "votes": null
    },
    {
      "id": "325167",
      "postDate": "05/08/2018 07:24:03",
      "content": "<p>Oh reading this makes me crying @Ahmet. Great out-of-the-box thinking! We knew that something was wrong with the order but we couldn't come up with anything good. My friend <a href=\"https://www.kaggle.com/mamasinkgs\">mamas</a> always said who finds the correct order of clicks easily wins this competition. He wasn't wrong about that... Too bad your team did not finish in price, you would have deserved it!!! </p>",
      "rawMarkdown": "Oh reading this makes me crying @Ahmet. Great out-of-the-box thinking! We knew that something was wrong with the order but we couldn't come up with anything good. My friend [mamas][1] always said who finds the correct order of clicks easily wins this competition. He wasn't wrong about that... Too bad your team did not finish in price, you would have deserved it!!! \n\n\n  [1]: https://www.kaggle.com/mamasinkgs",
      "votes": null
    },
    {
      "id": "325206",
      "postDate": "05/08/2018 08:13:39",
      "content": "<p>Btw I also tried IP%27. It did not work ;)</p>",
      "rawMarkdown": "Btw I also tried IP%27. It did not work ;)",
      "votes": null
    },
    {
      "id": "325232",
      "postDate": "05/08/2018 08:44:28",
      "content": "<p>I reached the same conclusion, after rereading <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677#321238\">my comment</a> that target seems ordered in test but not in train.  Ordering it in train was a very effective fix indeed.</p>",
      "rawMarkdown": "I reached the same conclusion, after rereading [my comment][1] that target seems ordered in test but not in train.  Ordering it in train was a very effective fix indeed.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677#321238",
      "votes": null
    },
    {
      "id": "325290",
      "postDate": "05/08/2018 09:24:55",
      "content": "<p>I wish I had seen before the end of the competition. As a newbie, I was overwhelmed by the amount of information and could not read everything. Very valuable. I am upvoting for you!</p>",
      "rawMarkdown": "I wish I had seen before the end of the competition. As a newbie, I was overwhelmed by the amount of information and could not read everything. Very valuable. I am upvoting for you!",
      "votes": null
    },
    {
      "id": "325319",
      "postDate": "05/08/2018 09:54:41",
      "content": "<p>Congrats @Ahmet and thanks for sharing</p>",
      "rawMarkdown": "Congrats @Ahmet and thanks for sharing",
      "votes": null
    },
    {
      "id": "325325",
      "postDate": "05/08/2018 10:07:20",
      "content": "<p>In fact, we did not find a way to solve duplicate data. Your method is perfect！！</p>",
      "rawMarkdown": "In fact, we did not find a way to solve duplicate data. Your method is perfect！！",
      "votes": null
    },
    {
      "id": "325399",
      "postDate": "05/08/2018 11:31:19",
      "content": "<p>Amazing. <br> Congrats :) </p>",
      "rawMarkdown": "Amazing. <br> Congrats :)",
      "votes": null
    },
    {
      "id": "325430",
      "postDate": "05/08/2018 12:04:32",
      "content": "<p>Forgot most important: congrats on the result, and for having found how to deal with duplicates before others!</p>",
      "rawMarkdown": "Forgot most important: congrats on the result, and for having found how to deal with duplicates before others!",
      "votes": null
    },
    {
      "id": "325432",
      "postDate": "05/08/2018 12:05:45",
      "content": "<p><a href=\"/plantsgo\">@plantsgo</a>, I was surprised to not see your score jump more after you shared it.  This makes your end result even more impressive, congrats!</p>",
      "rawMarkdown": "plantsgo, I was surprised to not see your score jump more after you shared it.  This makes your end result even more impressive, congrats!",
      "votes": null
    },
    {
      "id": "325497",
      "postDate": "05/08/2018 13:26:30",
      "content": "<p>Brave assumption on two tables and the good result is well deserved!</p>",
      "rawMarkdown": "Brave assumption on two tables and the good result is well deserved!",
      "votes": null
    },
    {
      "id": "325558",
      "postDate": "05/08/2018 14:45:12",
      "content": "<p>Now I'm looking forward to the first place not using this leak, otherwise I will regret why I didn't think of this method, ha ha ha...</p>",
      "rawMarkdown": "Now I'm looking forward to the first place not using this leak, otherwise I will regret why I didn't think of this method, ha ha ha...",
      "votes": null
    },
    {
      "id": "325598",
      "postDate": "05/08/2018 16:02:39",
      "content": "<p>Brilliant analysis. I spent a lot of time to think about this duplicate row issue, but still failed:(</p>",
      "rawMarkdown": "Brilliant analysis. I spent a lot of time to think about this duplicate row issue, but still failed:(",
      "votes": null
    },
    {
      "id": "325668",
      "postDate": "05/08/2018 17:27:52",
      "content": "<p>Thanks a lot for your positive comments:)</p>\n\n<p>@Danijel I would cry if IP%27 would turn out to be useful. It could be a good karma for me:) Thanks and congrats btw, we were lucky that we just scored just a very tiny bit higher.</p>\n\n<p><a href=\"/plantsgo\">@plantsgo</a> I was expecting you guys were using the same, because you had a jump after your discussion post about duplicates. Then it must be a coincidence.</p>\n\n<p>@Laevatein Indeed. I also got some benefit from removing channel categorical feature.</p>",
      "rawMarkdown": "Thanks a lot for your positive comments:)\n\n@Danijel I would cry if IP%27 would turn out to be useful. It could be a good karma for me:) Thanks and congrats btw, we were lucky that we just scored just a very tiny bit higher.\n\n@plantsgo I was expecting you guys were using the same, because you had a jump after your discussion post about duplicates. Then it must be a coincidence.\n\n@Laevatein Indeed. I also got some benefit from removing channel categorical feature.",
      "votes": null
    },
    {
      "id": "325671",
      "postDate": "05/08/2018 17:39:27",
      "content": "<p>The jump is from blending...We thought you were blending too, so we didn't notice...:(</p>",
      "rawMarkdown": "The jump is from blending...We thought you were blending too, so we didn't notice...:(",
      "votes": null
    },
    {
      "id": "325768",
      "postDate": "05/08/2018 20:56:11",
      "content": "<p>I still don't understand what channel is tbh...</p>",
      "rawMarkdown": "I still don't understand what channel is tbh...",
      "votes": null
    },
    {
      "id": "325899",
      "postDate": "05/09/2018 03:06:18",
      "content": "<p>Brilliant analysis!!!!! <br>\nThat is why I like data analysis so much! There are so many amazing pattern people can find in the data.\nThis solution is shock to me. Thanks you!!</p>\n\n<p>My idea is that it just like a race (whether next click is same app or not).\nThen I find the difference of ip_nextclick (or other group nextclick) and ip_device_os_app_nextclick is a good feature.\nThese feature is important in my lightGBM model but it doesn't work for LB score improvement.\nMaybe these feature will work in the condition that your solution mention about.</p>",
      "rawMarkdown": "Brilliant analysis!!!!!  \nThat is why I like data analysis so much! There are so many amazing pattern people can find in the data.\nThis solution is shock to me. Thanks you!!\n\nMy idea is that it just like a race (whether next click is same app or not).\nThen I find the difference of ip_nextclick (or other group nextclick) and ip_device_os_app_nextclick is a good feature.\nThese feature is important in my lightGBM model but it doesn't work for LB score improvement.\nMaybe these feature will work in the condition that your solution mention about.",
      "votes": null
    },
    {
      "id": "327229",
      "postDate": "05/11/2018 04:17:57",
      "content": "<p>This method is really really really remarkable, what a powerful data insight! Thank you so much for sharing this! Really hope I could upvote more than just one time!</p>",
      "rawMarkdown": "This method is really really really remarkable, what a powerful data insight! Thank you so much for sharing this! Really hope I could upvote more than just one time!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325142,
      "author_name": "a45632",
      "author_url": "",
      "post_date": "05/08/2018 06:59:31",
      "content": "<p>Great analysis,\nThanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325152,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "05/08/2018 07:15:43",
      "content": "<p>Very impressive discovery! You guys are soooooooooo smart. Thanks for sharing. BTW, does it mean that <strong>channel</strong> is kind of misleading?</p>",
      "votes": null,
      "replies": [
        {
          "id": 325768,
          "author_name": "thepathofd",
          "author_url": "",
          "post_date": "05/08/2018 20:56:11",
          "content": "<p>I still don't understand what channel is tbh...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325167,
      "author_name": "danijelk",
      "author_url": "",
      "post_date": "05/08/2018 07:24:03",
      "content": "<p>Oh reading this makes me crying @Ahmet. Great out-of-the-box thinking! We knew that something was wrong with the order but we couldn't come up with anything good. My friend <a href=\"https://www.kaggle.com/mamasinkgs\">mamas</a> always said who finds the correct order of clicks easily wins this competition. He wasn't wrong about that... Too bad your team did not finish in price, you would have deserved it!!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325206,
      "author_name": "danijelk",
      "author_url": "",
      "post_date": "05/08/2018 08:13:39",
      "content": "<p>Btw I also tried IP%27. It did not work ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325232,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/08/2018 08:44:28",
      "content": "<p>I reached the same conclusion, after rereading <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677#321238\">my comment</a> that target seems ordered in test but not in train.  Ordering it in train was a very effective fix indeed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325290,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/08/2018 09:24:55",
      "content": "<p>I wish I had seen before the end of the competition. As a newbie, I was overwhelmed by the amount of information and could not read everything. Very valuable. I am upvoting for you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325319,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "05/08/2018 09:54:41",
      "content": "<p>Congrats @Ahmet and thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325325,
      "author_name": "plantsgo",
      "author_url": "",
      "post_date": "05/08/2018 10:07:20",
      "content": "<p>In fact, we did not find a way to solve duplicate data. Your method is perfect！！</p>",
      "votes": null,
      "replies": [
        {
          "id": 325432,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/08/2018 12:05:45",
          "content": "<p><a href=\"/plantsgo\">@plantsgo</a>, I was surprised to not see your score jump more after you shared it.  This makes your end result even more impressive, congrats!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325558,
          "author_name": "plantsgo",
          "author_url": "",
          "post_date": "05/08/2018 14:45:12",
          "content": "<p>Now I'm looking forward to the first place not using this leak, otherwise I will regret why I didn't think of this method, ha ha ha...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325399,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "05/08/2018 11:31:19",
      "content": "<p>Amazing. <br> Congrats :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325430,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/08/2018 12:04:32",
      "content": "<p>Forgot most important: congrats on the result, and for having found how to deal with duplicates before others!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325497,
      "author_name": "cczaixian",
      "author_url": "",
      "post_date": "05/08/2018 13:26:30",
      "content": "<p>Brave assumption on two tables and the good result is well deserved!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325598,
      "author_name": "jyuan1986",
      "author_url": "",
      "post_date": "05/08/2018 16:02:39",
      "content": "<p>Brilliant analysis. I spent a lot of time to think about this duplicate row issue, but still failed:(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325668,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "05/08/2018 17:27:52",
      "content": "<p>Thanks a lot for your positive comments:)</p>\n\n<p>@Danijel I would cry if IP%27 would turn out to be useful. It could be a good karma for me:) Thanks and congrats btw, we were lucky that we just scored just a very tiny bit higher.</p>\n\n<p><a href=\"/plantsgo\">@plantsgo</a> I was expecting you guys were using the same, because you had a jump after your discussion post about duplicates. Then it must be a coincidence.</p>\n\n<p>@Laevatein Indeed. I also got some benefit from removing channel categorical feature.</p>",
      "votes": null,
      "replies": [
        {
          "id": 325671,
          "author_name": "plantsgo",
          "author_url": "",
          "post_date": "05/08/2018 17:39:27",
          "content": "<p>The jump is from blending...We thought you were blending too, so we didn't notice...:(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325899,
      "author_name": "yentianbao",
      "author_url": "",
      "post_date": "05/09/2018 03:06:18",
      "content": "<p>Brilliant analysis!!!!! <br>\nThat is why I like data analysis so much! There are so many amazing pattern people can find in the data.\nThis solution is shock to me. Thanks you!!</p>\n\n<p>My idea is that it just like a race (whether next click is same app or not).\nThen I find the difference of ip_nextclick (or other group nextclick) and ip_device_os_app_nextclick is a good feature.\nThese feature is important in my lightGBM model but it doesn't work for LB score improvement.\nMaybe these feature will work in the condition that your solution mention about.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327229,
      "author_name": "bangdasun",
      "author_url": "",
      "post_date": "05/11/2018 04:17:57",
      "content": "<p>This method is really really really remarkable, what a powerful data insight! Thank you so much for sharing this! Really hope I could upvote more than just one time!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325119": "As you have all noticed, there were duplicate samples with different labels. I will chronologically explain how I found a solution for it. \n\nIn the middle of competition, I think I have found time_to_next_click on [ip, app, device, os] feature earlier than public kernels and most of the LB. There was a discussion about the fact that click ids are not sorted. @Bojan stated that click ids are leaky in some TalkingData competitions. At that time, I didn't take this comment serious, and I even made a joke about IP%27. Maybe, some people had believed due to my high score which is actually thanks to relatively early discovery of time_to_next_click feature but not leak.\n\nThen I have put serious thinking effort on two things: Why time_to_next_click works best in [ip, app, device, os] group but nothing else like channel? Why there are duplicates with different labels?\n\nMy theory is TalkingData has 2 tables: Clicks table C(click_id, ip, app, device, os, channel, time) and Downloads table D(ip, app, device, os, time). So, the labels we see are not necessarily real downloads. This is how TalkingData maps D to C. If you check download times, some of the apps are downloaded 12 hours after they are clicked. This is not realistic. So a download in D is mapped to **the latest** click on C which has the same **(ip, app, device, os)** . **channel** is not important because there is no such thing as channel of a download. Then I noticed that we don't know which click is the latest if they are in the same second. I thought I can try ordering them by click_id, but this information is not available in training set. Then I thought if my logic is correct I can order the training set by (time, is_attributed) if and only if I have one time delta feature which is on (ip, app, device, os).\n\nSo, sorting the training data by **(time, is_attributed)**, and the test data by **(time, click_id)** gives +0.0005 boost on public LB if the time delta feature is used. It doesn't work with other time deltas because sorting by (time, is_attributed) is not valid for them and leak wrong information.\n\nMaybe others have better solution for this, but for me it was a big discovery. Since it was not a random discovery but serious thinking effort, I couldn't feel comfortable to share this during the competition.",
    "325142": "Great analysis,\nThanks for sharing.",
    "325152": "Very impressive discovery! You guys are soooooooooo smart. Thanks for sharing. BTW, does it mean that **channel** is kind of misleading?",
    "325167": "Oh reading this makes me crying @Ahmet. Great out-of-the-box thinking! We knew that something was wrong with the order but we couldn't come up with anything good. My friend [mamas][1] always said who finds the correct order of clicks easily wins this competition. He wasn't wrong about that... Too bad your team did not finish in price, you would have deserved it!!! \n\n\n  [1]: https://www.kaggle.com/mamasinkgs",
    "325206": "Btw I also tried IP%27. It did not work ;)",
    "325232": "I reached the same conclusion, after rereading [my comment][1] that target seems ordered in test but not in train.  Ordering it in train was a very effective fix indeed.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/55677#321238",
    "325290": "I wish I had seen before the end of the competition. As a newbie, I was overwhelmed by the amount of information and could not read everything. Very valuable. I am upvoting for you!",
    "325319": "Congrats @Ahmet and thanks for sharing",
    "325325": "In fact, we did not find a way to solve duplicate data. Your method is perfect！！",
    "325399": "Amazing. <br> Congrats :)",
    "325430": "Forgot most important: congrats on the result, and for having found how to deal with duplicates before others!",
    "325432": "plantsgo, I was surprised to not see your score jump more after you shared it.  This makes your end result even more impressive, congrats!",
    "325497": "Brave assumption on two tables and the good result is well deserved!",
    "325558": "Now I'm looking forward to the first place not using this leak, otherwise I will regret why I didn't think of this method, ha ha ha...",
    "325598": "Brilliant analysis. I spent a lot of time to think about this duplicate row issue, but still failed:(",
    "325668": "Thanks a lot for your positive comments:)\n\n@Danijel I would cry if IP%27 would turn out to be useful. It could be a good karma for me:) Thanks and congrats btw, we were lucky that we just scored just a very tiny bit higher.\n\n@plantsgo I was expecting you guys were using the same, because you had a jump after your discussion post about duplicates. Then it must be a coincidence.\n\n@Laevatein Indeed. I also got some benefit from removing channel categorical feature.",
    "325671": "The jump is from blending...We thought you were blending too, so we didn't notice...:(",
    "325768": "I still don't understand what channel is tbh...",
    "325899": "Brilliant analysis!!!!!  \nThat is why I like data analysis so much! There are so many amazing pattern people can find in the data.\nThis solution is shock to me. Thanks you!!\n\nMy idea is that it just like a race (whether next click is same app or not).\nThen I find the difference of ip_nextclick (or other group nextclick) and ip_device_os_app_nextclick is a good feature.\nThese feature is important in my lightGBM model but it doesn't work for LB score improvement.\nMaybe these feature will work in the condition that your solution mention about.",
    "327229": "This method is really really really remarkable, what a powerful data insight! Thank you so much for sharing this! Really hope I could upvote more than just one time!"
  },
  "source": "meta"
}