{
  "id": 56319,
  "title": "5th place story (pocket)",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56319",
  "author_name": "",
  "post_date": "2018-05-08T14:21:03.525373100Z",
  "votes": 51,
  "comment_count": 4,
  "views": 0,
  "content": "<p>So it was on a rainy day, one month till the end of the competition, that me and mamas met. <br>\nI was 6th publicLB, and mamas was 2nd. Mamas offered me a team merge, and I agreed. <br>\nWe had a nice (cheap) dinner at a restaurant in Tokyo. <br>\nWe started talking about our model, and yes, we couldn't stop talking. <br>\nWe talked so long that I nearly missed my train. (Who said this competition was named \"talking\"?)  </p>\n\n<p>Then after the day after we teamed up, Michael and Danijel joined. Two of the most talented kagglers. <br>\nWe thought that this team-up would alleviate the weak point our team had, namely NN and experience. <br>\nAs you may know, Michael has a legendary win with his NN on the Porto Seguro competition, and <br>\nDanijel has a lot of experience with only one gold medal left for his GrandMaster title. <br>\nHappy with the team composition, we aimed high. Higher than I could have ever imagined. <br>\nThe top prize.  </p>\n\n<p>Here is what I tried.  </p>\n\n<p>When we met, mamas told me that he had a few hundred features(!), and was training on a 1.4TB RAM GCP machine (WTF!) <br>\nI instantly understood that this competition was not a competition for my single model to shine. <br>\nSo I decided to dig really hard on exotic techniques. <br>\nThe two most important ones I tried (which failed) were <br>\nA)Pseudo Labeling B)Data Augmentation  </p>\n\n<p>A)Pseudo Labeling <br>\nDid anyone succeed with this? I tried PL with lightgbm in many ways. <br>\nDifferent types of objectives(binary, regression, xentropy), different amount of data to mix, soft labels, hard labels. <br>\nNothing exceeded the score I could achieve with a single run. <br>\nRegression exceeded the first round score, but still lower than the single run with binary objective.  </p>\n\n<p>B)Data Augmentation <br>\nI tried some form of pseudo data augmentation used in the Recruit competition. <br>\n<a href=\"https://www.kaggle.com/pureheart/1st-place-lgb-model-public-0-470-private-0-502\">https://www.kaggle.com/pureheart/1st-place-lgb-model-public-0-470-private-0-502</a> <br>\nFor example, I divided the data into 4hour windows, and slid the window by 2 hours to make 2x data. <br>\nex) 0-4am, 2-6am, 4-8am... and so on, and I extracted the features from each of the window. <br>\nThis did not work either.  </p>\n\n<p>So there I was, only with two weeks left till the end of the competition. <br>\nI realized that I could not make a difference with modeling techniques, and started to focus on my single model. Since by then, mamas had a 0.9826 model. I knew I should make a simple, but diverse model that adds to the blend. <br>\nWhat did that look like?  </p>\n\n<p>C) Single model(0.9828private) <br>\nSince mamas had a gigantic model, I focused on simplicity. <br>\nI had roughly 30 features(including the original ones). <br>\nThe most important ones (if you have not noticed) are, 1)nextClicks 2)ip-nunique features <br>\nI extracted every feature day-by-day. This was unlike mamas or Danijel's model, so I think it added diversity. <br>\nThe model was seed-averaged with 5 seeds for the final submission. (+0.0001) <br>\nTrained on a GCP machine. 64CoreCPU, 416GB RAM. <br>\nI could have used a smaller machine by coding more smartly, but for me, time(and score) was important than money.  </p>\n\n<p>In the end, we threw in our models into Michael's blending NN. <br>\n0.9836(mamas), 0.9828(pocket), 0.9827(Danijel), and a few NN from Michael and Danijel. <br>\nOut comes 0.9840, 5th.  </p>\n\n<p>We couldn't get the top prize, but at the same time, I am happy with my gold medal. <br>\nHappy Kaggling, and thank you for reading.  </p>\n\n<p>P.S. You will be hearing from other Superstars from our team, in a few days...</p>",
  "messages": [
    {
      "id": "325540",
      "postDate": "05/08/2018 14:21:03",
      "content": "<p>So it was on a rainy day, one month till the end of the competition, that me and mamas met. <br>\nI was 6th publicLB, and mamas was 2nd. Mamas offered me a team merge, and I agreed. <br>\nWe had a nice (cheap) dinner at a restaurant in Tokyo. <br>\nWe started talking about our model, and yes, we couldn't stop talking. <br>\nWe talked so long that I nearly missed my train. (Who said this competition was named \"talking\"?)  </p>\n\n<p>Then after the day after we teamed up, Michael and Danijel joined. Two of the most talented kagglers. <br>\nWe thought that this team-up would alleviate the weak point our team had, namely NN and experience. <br>\nAs you may know, Michael has a legendary win with his NN on the Porto Seguro competition, and <br>\nDanijel has a lot of experience with only one gold medal left for his GrandMaster title. <br>\nHappy with the team composition, we aimed high. Higher than I could have ever imagined. <br>\nThe top prize.  </p>\n\n<p>Here is what I tried.  </p>\n\n<p>When we met, mamas told me that he had a few hundred features(!), and was training on a 1.4TB RAM GCP machine (WTF!) <br>\nI instantly understood that this competition was not a competition for my single model to shine. <br>\nSo I decided to dig really hard on exotic techniques. <br>\nThe two most important ones I tried (which failed) were <br>\nA)Pseudo Labeling B)Data Augmentation  </p>\n\n<p>A)Pseudo Labeling <br>\nDid anyone succeed with this? I tried PL with lightgbm in many ways. <br>\nDifferent types of objectives(binary, regression, xentropy), different amount of data to mix, soft labels, hard labels. <br>\nNothing exceeded the score I could achieve with a single run. <br>\nRegression exceeded the first round score, but still lower than the single run with binary objective.  </p>\n\n<p>B)Data Augmentation <br>\nI tried some form of pseudo data augmentation used in the Recruit competition. <br>\n<a href=\"https://www.kaggle.com/pureheart/1st-place-lgb-model-public-0-470-private-0-502\">https://www.kaggle.com/pureheart/1st-place-lgb-model-public-0-470-private-0-502</a> <br>\nFor example, I divided the data into 4hour windows, and slid the window by 2 hours to make 2x data. <br>\nex) 0-4am, 2-6am, 4-8am... and so on, and I extracted the features from each of the window. <br>\nThis did not work either.  </p>\n\n<p>So there I was, only with two weeks left till the end of the competition. <br>\nI realized that I could not make a difference with modeling techniques, and started to focus on my single model. Since by then, mamas had a 0.9826 model. I knew I should make a simple, but diverse model that adds to the blend. <br>\nWhat did that look like?  </p>\n\n<p>C) Single model(0.9828private) <br>\nSince mamas had a gigantic model, I focused on simplicity. <br>\nI had roughly 30 features(including the original ones). <br>\nThe most important ones (if you have not noticed) are, 1)nextClicks 2)ip-nunique features <br>\nI extracted every feature day-by-day. This was unlike mamas or Danijel's model, so I think it added diversity. <br>\nThe model was seed-averaged with 5 seeds for the final submission. (+0.0001) <br>\nTrained on a GCP machine. 64CoreCPU, 416GB RAM. <br>\nI could have used a smaller machine by coding more smartly, but for me, time(and score) was important than money.  </p>\n\n<p>In the end, we threw in our models into Michael's blending NN. <br>\n0.9836(mamas), 0.9828(pocket), 0.9827(Danijel), and a few NN from Michael and Danijel. <br>\nOut comes 0.9840, 5th.  </p>\n\n<p>We couldn't get the top prize, but at the same time, I am happy with my gold medal. <br>\nHappy Kaggling, and thank you for reading.  </p>\n\n<p>P.S. You will be hearing from other Superstars from our team, in a few days...</p>",
      "rawMarkdown": "So it was on a rainy day, one month till the end of the competition, that me and mamas met.  \nI was 6th publicLB, and mamas was 2nd. Mamas offered me a team merge, and I agreed.  \nWe had a nice (cheap) dinner at a restaurant in Tokyo.  \nWe started talking about our model, and yes, we couldn't stop talking.   \nWe talked so long that I nearly missed my train. (Who said this competition was named \"talking\"?)  \n  \nThen after the day after we teamed up, Michael and Danijel joined. Two of the most talented kagglers.  \nWe thought that this team-up would alleviate the weak point our team had, namely NN and experience.  \nAs you may know, Michael has a legendary win with his NN on the Porto Seguro competition, and  \nDanijel has a lot of experience with only one gold medal left for his GrandMaster title.  \nHappy with the team composition, we aimed high. Higher than I could have ever imagined.  \nThe top prize.  \n  \nHere is what I tried.  \n  \nWhen we met, mamas told me that he had a few hundred features(!), and was training on a 1.4TB RAM GCP machine (WTF!)  \nI instantly understood that this competition was not a competition for my single model to shine.  \nSo I decided to dig really hard on exotic techniques.  \nThe two most important ones I tried (which failed) were  \nA)Pseudo Labeling B)Data Augmentation  \n  \nA)Pseudo Labeling  \nDid anyone succeed with this? I tried PL with lightgbm in many ways.  \nDifferent types of objectives(binary, regression, xentropy), different amount of data to mix, soft labels, hard labels.  \nNothing exceeded the score I could achieve with a single run.   \nRegression exceeded the first round score, but still lower than the single run with binary objective.  \n  \nB)Data Augmentation  \nI tried some form of pseudo data augmentation used in the Recruit competition.  \nhttps://www.kaggle.com/pureheart/1st-place-lgb-model-public-0-470-private-0-502  \nFor example, I divided the data into 4hour windows, and slid the window by 2 hours to make 2x data.  \nex) 0-4am, 2-6am, 4-8am... and so on, and I extracted the features from each of the window.  \nThis did not work either.  \n  \nSo there I was, only with two weeks left till the end of the competition.  \nI realized that I could not make a difference with modeling techniques, and started to focus on my single model. Since by then, mamas had a 0.9826 model. I knew I should make a simple, but diverse model that adds to the blend.  \nWhat did that look like?  \n  \nC) Single model(0.9828private)  \nSince mamas had a gigantic model, I focused on simplicity.  \nI had roughly 30 features(including the original ones).  \nThe most important ones (if you have not noticed) are, 1)nextClicks 2)ip-nunique features  \nI extracted every feature day-by-day. This was unlike mamas or Danijel's model, so I think it added diversity.  \nThe model was seed-averaged with 5 seeds for the final submission. (+0.0001)   \nTrained on a GCP machine. 64CoreCPU, 416GB RAM.  \nI could have used a smaller machine by coding more smartly, but for me, time(and score) was important than money.  \n  \nIn the end, we threw in our models into Michael's blending NN.  \n0.9836(mamas), 0.9828(pocket), 0.9827(Danijel), and a few NN from Michael and Danijel.  \nOut comes 0.9840, 5th.  \n  \nWe couldn't get the top prize, but at the same time, I am happy with my gold medal.  \nHappy Kaggling, and thank you for reading.  \n\nP.S. You will be hearing from other Superstars from our team, in a few days...",
      "votes": null
    },
    {
      "id": "325546",
      "postDate": "05/08/2018 14:30:52",
      "content": "<p>Thanks for sharing, great story.  And congrat for the result!</p>",
      "rawMarkdown": "Thanks for sharing, great story.  And congrat for the result!",
      "votes": null
    },
    {
      "id": "326081",
      "postDate": "05/09/2018 08:39:56",
      "content": "<p>As a newbie, I have been considering about adding more RAM to my pc.But after reading this, I give up. Maybe saving money for the cloud machine is a much better choice:)\nAnd, thanks for sharing.</p>",
      "rawMarkdown": "As a newbie, I have been considering about adding more RAM to my pc.But after reading this, I give up. Maybe saving money for the cloud machine is a much better choice:)\nAnd, thanks for sharing.",
      "votes": null
    },
    {
      "id": "326189",
      "postDate": "05/09/2018 12:33:22",
      "content": "<p>My local machine has 16GB RAM and it was enough to win a gold medal in the Recruit competition. <br>\n<a href=\"https://www.kaggle.com/c/recruit-restaurant-visitor-forecasting\">https://www.kaggle.com/c/recruit-restaurant-visitor-forecasting</a> <br>\nSo I'd say, keep your hopes high, and keep on trying :)  </p>\n\n<p>btw you can get a free $300 credit for GCP, which covered most of the cost for this competition. <br>\n(Although I think Kaggle should host more competitions with small data size....)</p>",
      "rawMarkdown": "My local machine has 16GB RAM and it was enough to win a gold medal in the Recruit competition.   \nhttps://www.kaggle.com/c/recruit-restaurant-visitor-forecasting  \nSo I'd say, keep your hopes high, and keep on trying :)  \n\nbtw you can get a free $300 credit for GCP, which covered most of the cost for this competition.  \n(Although I think Kaggle should host more competitions with small data size....)",
      "votes": null
    },
    {
      "id": "326575",
      "postDate": "05/10/2018 01:02:53",
      "content": "<p>I'm happy to know that somebody can win a gold medal with a 16GB RAM machine.And certainly, I'll check GCP out.Thanks!</p>",
      "rawMarkdown": "I'm happy to know that somebody can win a gold medal with a 16GB RAM machine.And certainly, I'll check GCP out.Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325546,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/08/2018 14:30:52",
      "content": "<p>Thanks for sharing, great story.  And congrat for the result!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326081,
      "author_name": "xiguapi",
      "author_url": "",
      "post_date": "05/09/2018 08:39:56",
      "content": "<p>As a newbie, I have been considering about adding more RAM to my pc.But after reading this, I give up. Maybe saving money for the cloud machine is a much better choice:)\nAnd, thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 326189,
          "author_name": "pocketsuteado",
          "author_url": "",
          "post_date": "05/09/2018 12:33:22",
          "content": "<p>My local machine has 16GB RAM and it was enough to win a gold medal in the Recruit competition. <br>\n<a href=\"https://www.kaggle.com/c/recruit-restaurant-visitor-forecasting\">https://www.kaggle.com/c/recruit-restaurant-visitor-forecasting</a> <br>\nSo I'd say, keep your hopes high, and keep on trying :)  </p>\n\n<p>btw you can get a free $300 credit for GCP, which covered most of the cost for this competition. <br>\n(Although I think Kaggle should host more competitions with small data size....)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 326575,
          "author_name": "xiguapi",
          "author_url": "",
          "post_date": "05/10/2018 01:02:53",
          "content": "<p>I'm happy to know that somebody can win a gold medal with a 16GB RAM machine.And certainly, I'll check GCP out.Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "325540": "So it was on a rainy day, one month till the end of the competition, that me and mamas met.  \nI was 6th publicLB, and mamas was 2nd. Mamas offered me a team merge, and I agreed.  \nWe had a nice (cheap) dinner at a restaurant in Tokyo.  \nWe started talking about our model, and yes, we couldn't stop talking.   \nWe talked so long that I nearly missed my train. (Who said this competition was named \"talking\"?)  \n  \nThen after the day after we teamed up, Michael and Danijel joined. Two of the most talented kagglers.  \nWe thought that this team-up would alleviate the weak point our team had, namely NN and experience.  \nAs you may know, Michael has a legendary win with his NN on the Porto Seguro competition, and  \nDanijel has a lot of experience with only one gold medal left for his GrandMaster title.  \nHappy with the team composition, we aimed high. Higher than I could have ever imagined.  \nThe top prize.  \n  \nHere is what I tried.  \n  \nWhen we met, mamas told me that he had a few hundred features(!), and was training on a 1.4TB RAM GCP machine (WTF!)  \nI instantly understood that this competition was not a competition for my single model to shine.  \nSo I decided to dig really hard on exotic techniques.  \nThe two most important ones I tried (which failed) were  \nA)Pseudo Labeling B)Data Augmentation  \n  \nA)Pseudo Labeling  \nDid anyone succeed with this? I tried PL with lightgbm in many ways.  \nDifferent types of objectives(binary, regression, xentropy), different amount of data to mix, soft labels, hard labels.  \nNothing exceeded the score I could achieve with a single run.   \nRegression exceeded the first round score, but still lower than the single run with binary objective.  \n  \nB)Data Augmentation  \nI tried some form of pseudo data augmentation used in the Recruit competition.  \nhttps://www.kaggle.com/pureheart/1st-place-lgb-model-public-0-470-private-0-502  \nFor example, I divided the data into 4hour windows, and slid the window by 2 hours to make 2x data.  \nex) 0-4am, 2-6am, 4-8am... and so on, and I extracted the features from each of the window.  \nThis did not work either.  \n  \nSo there I was, only with two weeks left till the end of the competition.  \nI realized that I could not make a difference with modeling techniques, and started to focus on my single model. Since by then, mamas had a 0.9826 model. I knew I should make a simple, but diverse model that adds to the blend.  \nWhat did that look like?  \n  \nC) Single model(0.9828private)  \nSince mamas had a gigantic model, I focused on simplicity.  \nI had roughly 30 features(including the original ones).  \nThe most important ones (if you have not noticed) are, 1)nextClicks 2)ip-nunique features  \nI extracted every feature day-by-day. This was unlike mamas or Danijel's model, so I think it added diversity.  \nThe model was seed-averaged with 5 seeds for the final submission. (+0.0001)   \nTrained on a GCP machine. 64CoreCPU, 416GB RAM.  \nI could have used a smaller machine by coding more smartly, but for me, time(and score) was important than money.  \n  \nIn the end, we threw in our models into Michael's blending NN.  \n0.9836(mamas), 0.9828(pocket), 0.9827(Danijel), and a few NN from Michael and Danijel.  \nOut comes 0.9840, 5th.  \n  \nWe couldn't get the top prize, but at the same time, I am happy with my gold medal.  \nHappy Kaggling, and thank you for reading.  \n\nP.S. You will be hearing from other Superstars from our team, in a few days...",
    "325546": "Thanks for sharing, great story.  And congrat for the result!",
    "326081": "As a newbie, I have been considering about adding more RAM to my pc.But after reading this, I give up. Maybe saving money for the cloud machine is a much better choice:)\nAnd, thanks for sharing.",
    "326189": "My local machine has 16GB RAM and it was enough to win a gold medal in the Recruit competition.   \nhttps://www.kaggle.com/c/recruit-restaurant-visitor-forecasting  \nSo I'd say, keep your hopes high, and keep on trying :)  \n\nbtw you can get a free $300 credit for GCP, which covered most of the cost for this competition.  \n(Although I think Kaggle should host more competitions with small data size....)",
    "326575": "I'm happy to know that somebody can win a gold medal with a 16GB RAM machine.And certainly, I'll check GCP out.Thanks!"
  },
  "source": "meta"
}