{
  "id": 56317,
  "title": "[late sub 0.9828771]How Boosting from 0.9671 to 0.9823 my first competition on kaggle",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/definitely-maybe-late-sub-0-9828771-how-boosting-f",
  "author_name": "",
  "post_date": "2018-05-11T15:22:58.243Z",
  "votes": 21,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, THANKS to kaggle and talking data to launch this great competition. Congrats to all the winners and participants who contributed to this game, I really learned so much from kernels and discussions from you guys. Awesome.</p>\n\n<p>For my work, ranked 65th, not that good enough but I'd like to share what i've been done throughout this game. I kept track of everyday works like the following.</p>\n\n<p><strong>1%</strong> <br>\nAt the very beginning I had no idea with the ip, app...so I brutely made LOTS OF groupby counts features (~60). Ran a lightgbm on day9, to reach 0.9671.</p>\n\n<p><strong>2%</strong> <br>\nAdded LOTS OF conversion-rate features (~20) to reach 0.9681.</p>\n\n<p><strong>50%</strong> <br>\nAt this point I learned from a few kernels that time-related features are keys to this competition. Added a few time-delta features jumping to 0.9799.</p>\n\n<p><strong>75%</strong> <br>\nRead discussions every morning and found test_supplement.csv. Run all features with that again to reach 0.9807.</p>\n\n<p><strong>90%</strong> <br>\nUse full train.csv to train lightgbm with the same features to reach 0.9810.</p>\n\n<p><strong>95%</strong> <br>\nUse hours 4,5,9,10,13,14 as validset and train again to reach 0.9813.</p>\n\n<p><strong>98%</strong> <br>\nAdd two more time-till-next-click groupby, to reach 0.9815. (At the very last day...)</p>\n\n<p><strong>99%</strong> <br>\nSimple post blend 3 or 4 lightgbms to finally reach 0.9816 public score. And my ranking for private leaderboard is 1 position upper (65th) than public leaderboard (66th).</p>\n\n<p>My final model is a single lightgbm with around 40 features. One interesting finding is that my local valid score is around 0.9823 which exactly match the private LB score. (public is 0.9816 though)</p>\n\n<hr>\n\n<p><strong>WHAT didn't work in my model</strong> <br>\n 1. empirical bayes target encoding for such as ip-app-download-rate didn't improve anymore. <br>\n 2. click in the next/last hour/24hour didn't work... <br>\n 3. feature that indicates ip's coming day...since I learned that newly coming ip has higher CVR...but didn't work... <br>\n 4. sample weighed for testhour in training lightgbm didn't work...</p>\n\n<p><strong>WHAT i wish i should've done</strong> <br>\n 1. should've done more time-delta features, as I found in the last day...Due to my ugly coding, I felt  lazy to add more of that... <br>\n 2. more cleverly dealing with duplicates (ip, app, device, os, channel, click_time). I came to <a href=\"/plantsgo\">@plantsgo</a>'s discussion on modifying duplicates and @CPMP 's deep thoughts on that...Intuitively I agree that the last line of duplicates should be a 1, if I just have more time to reorder and try that data training again to see the LB score.</p>\n\n<p>Anyway, really learned a lot from this competition, and really willing to join the next one... and see u guys there ;)</p>\n\n<hr>\n\n<p><strong>Updated 5/11 <br>\nI did reorder duplicates clicks by (time,is_attributed), and added next 2,3 clicks time-delta features. SO that gave me 0.9828771 on private leaderboard which could jump to 21th place.</strong></p>",
  "messages": [
    {
      "id": "325531",
      "postDate": "05/08/2018 14:06:17",
      "content": "<p>First of all, THANKS to kaggle and talking data to launch this great competition. Congrats to all the winners and participants who contributed to this game, I really learned so much from kernels and discussions from you guys. Awesome.</p>\n\n<p>For my work, ranked 65th, not that good enough but I'd like to share what i've been done throughout this game. I kept track of everyday works like the following.</p>\n\n<p><strong>1%</strong> <br>\nAt the very beginning I had no idea with the ip, app...so I brutely made LOTS OF groupby counts features (~60). Ran a lightgbm on day9, to reach 0.9671.</p>\n\n<p><strong>2%</strong> <br>\nAdded LOTS OF conversion-rate features (~20) to reach 0.9681.</p>\n\n<p><strong>50%</strong> <br>\nAt this point I learned from a few kernels that time-related features are keys to this competition. Added a few time-delta features jumping to 0.9799.</p>\n\n<p><strong>75%</strong> <br>\nRead discussions every morning and found test_supplement.csv. Run all features with that again to reach 0.9807.</p>\n\n<p><strong>90%</strong> <br>\nUse full train.csv to train lightgbm with the same features to reach 0.9810.</p>\n\n<p><strong>95%</strong> <br>\nUse hours 4,5,9,10,13,14 as validset and train again to reach 0.9813.</p>\n\n<p><strong>98%</strong> <br>\nAdd two more time-till-next-click groupby, to reach 0.9815. (At the very last day...)</p>\n\n<p><strong>99%</strong> <br>\nSimple post blend 3 or 4 lightgbms to finally reach 0.9816 public score. And my ranking for private leaderboard is 1 position upper (65th) than public leaderboard (66th).</p>\n\n<p>My final model is a single lightgbm with around 40 features. One interesting finding is that my local valid score is around 0.9823 which exactly match the private LB score. (public is 0.9816 though)</p>\n\n<hr>\n\n<p><strong>WHAT didn't work in my model</strong> <br>\n 1. empirical bayes target encoding for such as ip-app-download-rate didn't improve anymore. <br>\n 2. click in the next/last hour/24hour didn't work... <br>\n 3. feature that indicates ip's coming day...since I learned that newly coming ip has higher CVR...but didn't work... <br>\n 4. sample weighed for testhour in training lightgbm didn't work...</p>\n\n<p><strong>WHAT i wish i should've done</strong> <br>\n 1. should've done more time-delta features, as I found in the last day...Due to my ugly coding, I felt  lazy to add more of that... <br>\n 2. more cleverly dealing with duplicates (ip, app, device, os, channel, click_time). I came to <a href=\"/plantsgo\">@plantsgo</a>'s discussion on modifying duplicates and @CPMP 's deep thoughts on that...Intuitively I agree that the last line of duplicates should be a 1, if I just have more time to reorder and try that data training again to see the LB score.</p>\n\n<p>Anyway, really learned a lot from this competition, and really willing to join the next one... and see u guys there ;)</p>\n\n<hr>\n\n<p><strong>Updated 5/11 <br>\nI did reorder duplicates clicks by (time,is_attributed), and added next 2,3 clicks time-delta features. SO that gave me 0.9828771 on private leaderboard which could jump to 21th place.</strong></p>",
      "rawMarkdown": "First of all, THANKS to kaggle and talking data to launch this great competition. Congrats to all the winners and participants who contributed to this game, I really learned so much from kernels and discussions from you guys. Awesome.\n\nFor my work, ranked 65th, not that good enough but I'd like to share what i've been done throughout this game. I kept track of everyday works like the following.\n\n**1%**  \nAt the very beginning I had no idea with the ip, app...so I brutely made LOTS OF groupby counts features (~60). Ran a lightgbm on day9, to reach 0.9671.\n\n**2%**  \nAdded LOTS OF conversion-rate features (~20) to reach 0.9681.\n\n**50%**  \nAt this point I learned from a few kernels that time-related features are keys to this competition. Added a few time-delta features jumping to 0.9799.\n\n**75%**  \nRead discussions every morning and found test_supplement.csv. Run all features with that again to reach 0.9807.\n\n**90%**  \nUse full train.csv to train lightgbm with the same features to reach 0.9810.\n\n**95%**  \nUse hours 4,5,9,10,13,14 as validset and train again to reach 0.9813.\n\n**98%**  \nAdd two more time-till-next-click groupby, to reach 0.9815. (At the very last day...)\n\n**99%**  \nSimple post blend 3 or 4 lightgbms to finally reach 0.9816 public score. And my ranking for private leaderboard is 1 position upper (65th) than public leaderboard (66th).\n\nMy final model is a single lightgbm with around 40 features. One interesting finding is that my local valid score is around 0.9823 which exactly match the private LB score. (public is 0.9816 though)\n\n----------\n**WHAT didn't work in my model**  \n 1. empirical bayes target encoding for such as ip-app-download-rate didn't improve anymore.  \n 2. click in the next/last hour/24hour didn't work...  \n 3. feature that indicates ip's coming day...since I learned that newly coming ip has higher CVR...but didn't work...  \n 4. sample weighed for testhour in training lightgbm didn't work...\n\n\n**WHAT i wish i should've done**  \n 1. should've done more time-delta features, as I found in the last day...Due to my ugly coding, I felt  lazy to add more of that...  \n 2. more cleverly dealing with duplicates (ip, app, device, os, channel, click_time). I came to @plantsgo's discussion on modifying duplicates and @CPMP 's deep thoughts on that...Intuitively I agree that the last line of duplicates should be a 1, if I just have more time to reorder and try that data training again to see the LB score.\n\n\nAnyway, really learned a lot from this competition, and really willing to join the next one... and see u guys there ;)\n\n\n----------\n**Updated 5/11  \nI did reorder duplicates clicks by (time,is_attributed), and added next 2,3 clicks time-delta features. SO that gave me 0.9828771 on private leaderboard which could jump to 21th place.**",
      "votes": null
    },
    {
      "id": "325563",
      "postDate": "05/08/2018 14:49:48",
      "content": "<p>Congrats and thanks for your share! I want to know how many RAM do you have ?How do you deal with big data size？</p>",
      "rawMarkdown": "Congrats and thanks for your share! I want to know how many RAM do you have ?How do you deal with big data size？",
      "votes": null
    },
    {
      "id": "325571",
      "postDate": "05/08/2018 15:04:29",
      "content": "<p>Thanks. I work on a 32 cores workstation with around 40GB ram. I use pandas to read and process dataset by chunks, and within each chunk I use multiprocessing to process \"subchunks\" and finally combine all outputs.</p>",
      "rawMarkdown": "Thanks. I work on a 32 cores workstation with around 40GB ram. I use pandas to read and process dataset by chunks, and within each chunk I use multiprocessing to process \"subchunks\" and finally combine all outputs.",
      "votes": null
    },
    {
      "id": "325948",
      "postDate": "05/09/2018 05:27:37",
      "content": "<p>Thanks for sharing. Look like you had a very good experience :)\nBTW, could you plz share how did you setup your local validation?</p>",
      "rawMarkdown": "Thanks for sharing. Look like you had a very good experience :)\nBTW, could you plz share how did you setup your local validation?",
      "votes": null
    },
    {
      "id": "325957",
      "postDate": "05/09/2018 05:43:22",
      "content": "<p>Thanks for sharing, How many RAM do you have?</p>",
      "rawMarkdown": "Thanks for sharing, How many RAM do you have?",
      "votes": null
    },
    {
      "id": "326292",
      "postDate": "05/09/2018 14:17:35",
      "content": "<p>Thank you. Weeks before the end, I was using last 10m rows of day9 as valid, that was always a .004 gap between local and public score. Until I came to set hour 4,5,9,10,13,14 as valid, I actually have three random split valid sets. As I tuned the lightgbm more and more, the gap finally narrowed down to .0005~.0008. (.9823 local average, .9816 public, .9823 private)</p>",
      "rawMarkdown": "Thank you. Weeks before the end, I was using last 10m rows of day9 as valid, that was always a .004 gap between local and public score. Until I came to set hour 4,5,9,10,13,14 as valid, I actually have three random split valid sets. As I tuned the lightgbm more and more, the gap finally narrowed down to .0005~.0008. (.9823 local average, .9816 public, .9823 private)",
      "votes": null
    },
    {
      "id": "326297",
      "postDate": "05/09/2018 14:23:46",
      "content": "<p>Thanks. I was using 40GB ram machine. I use lightgbm CLI api, and two_round_loading=true, that can load more than 40GB datasets to train.</p>",
      "rawMarkdown": "Thanks. I was using 40GB ram machine. I use lightgbm CLI api, and two_round_loading=true, that can load more than 40GB datasets to train.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325563,
      "author_name": "dean1977",
      "author_url": "",
      "post_date": "05/08/2018 14:49:48",
      "content": "<p>Congrats and thanks for your share! I want to know how many RAM do you have ?How do you deal with big data size？</p>",
      "votes": null,
      "replies": [
        {
          "id": 325571,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "05/08/2018 15:04:29",
          "content": "<p>Thanks. I work on a 32 cores workstation with around 40GB ram. I use pandas to read and process dataset by chunks, and within each chunk I use multiprocessing to process \"subchunks\" and finally combine all outputs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325948,
      "author_name": "nguyentp",
      "author_url": "",
      "post_date": "05/09/2018 05:27:37",
      "content": "<p>Thanks for sharing. Look like you had a very good experience :)\nBTW, could you plz share how did you setup your local validation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 326292,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "05/09/2018 14:17:35",
          "content": "<p>Thank you. Weeks before the end, I was using last 10m rows of day9 as valid, that was always a .004 gap between local and public score. Until I came to set hour 4,5,9,10,13,14 as valid, I actually have three random split valid sets. As I tuned the lightgbm more and more, the gap finally narrowed down to .0005~.0008. (.9823 local average, .9816 public, .9823 private)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325957,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "05/09/2018 05:43:22",
      "content": "<p>Thanks for sharing, How many RAM do you have?</p>",
      "votes": null,
      "replies": [
        {
          "id": 326297,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "05/09/2018 14:23:46",
          "content": "<p>Thanks. I was using 40GB ram machine. I use lightgbm CLI api, and two_round_loading=true, that can load more than 40GB datasets to train.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "325531": "First of all, THANKS to kaggle and talking data to launch this great competition. Congrats to all the winners and participants who contributed to this game, I really learned so much from kernels and discussions from you guys. Awesome.\n\nFor my work, ranked 65th, not that good enough but I'd like to share what i've been done throughout this game. I kept track of everyday works like the following.\n\n**1%**  \nAt the very beginning I had no idea with the ip, app...so I brutely made LOTS OF groupby counts features (~60). Ran a lightgbm on day9, to reach 0.9671.\n\n**2%**  \nAdded LOTS OF conversion-rate features (~20) to reach 0.9681.\n\n**50%**  \nAt this point I learned from a few kernels that time-related features are keys to this competition. Added a few time-delta features jumping to 0.9799.\n\n**75%**  \nRead discussions every morning and found test_supplement.csv. Run all features with that again to reach 0.9807.\n\n**90%**  \nUse full train.csv to train lightgbm with the same features to reach 0.9810.\n\n**95%**  \nUse hours 4,5,9,10,13,14 as validset and train again to reach 0.9813.\n\n**98%**  \nAdd two more time-till-next-click groupby, to reach 0.9815. (At the very last day...)\n\n**99%**  \nSimple post blend 3 or 4 lightgbms to finally reach 0.9816 public score. And my ranking for private leaderboard is 1 position upper (65th) than public leaderboard (66th).\n\nMy final model is a single lightgbm with around 40 features. One interesting finding is that my local valid score is around 0.9823 which exactly match the private LB score. (public is 0.9816 though)\n\n----------\n**WHAT didn't work in my model**  \n 1. empirical bayes target encoding for such as ip-app-download-rate didn't improve anymore.  \n 2. click in the next/last hour/24hour didn't work...  \n 3. feature that indicates ip's coming day...since I learned that newly coming ip has higher CVR...but didn't work...  \n 4. sample weighed for testhour in training lightgbm didn't work...\n\n\n**WHAT i wish i should've done**  \n 1. should've done more time-delta features, as I found in the last day...Due to my ugly coding, I felt  lazy to add more of that...  \n 2. more cleverly dealing with duplicates (ip, app, device, os, channel, click_time). I came to @plantsgo's discussion on modifying duplicates and @CPMP 's deep thoughts on that...Intuitively I agree that the last line of duplicates should be a 1, if I just have more time to reorder and try that data training again to see the LB score.\n\n\nAnyway, really learned a lot from this competition, and really willing to join the next one... and see u guys there ;)\n\n\n----------\n**Updated 5/11  \nI did reorder duplicates clicks by (time,is_attributed), and added next 2,3 clicks time-delta features. SO that gave me 0.9828771 on private leaderboard which could jump to 21th place.**",
    "325563": "Congrats and thanks for your share! I want to know how many RAM do you have ?How do you deal with big data size？",
    "325571": "Thanks. I work on a 32 cores workstation with around 40GB ram. I use pandas to read and process dataset by chunks, and within each chunk I use multiprocessing to process \"subchunks\" and finally combine all outputs.",
    "325948": "Thanks for sharing. Look like you had a very good experience :)\nBTW, could you plz share how did you setup your local validation?",
    "325957": "Thanks for sharing, How many RAM do you have?",
    "326292": "Thank you. Weeks before the end, I was using last 10m rows of day9 as valid, that was always a .004 gap between local and public score. Until I came to set hour 4,5,9,10,13,14 as valid, I actually have three random split valid sets. As I tuned the lightgbm more and more, the gap finally narrowed down to .0005~.0008. (.9823 local average, .9816 public, .9823 private)",
    "326297": "Thanks. I was using 40GB ram machine. I use lightgbm CLI api, and two_round_loading=true, that can load more than 40GB datasets to train."
  },
  "source": "meta"
}