{
  "id": 56404,
  "title": "191st Place Solution Or more like lessons learnt....",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56404",
  "author_name": "Setu Chokshi",
  "post_date": "2018-05-09T12:58:56.141000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This was a crazy competition with the release of last minute of 9811 kernel, I stayed up the night to retune my models and see if I can beat that score. I edged it out only slightly, but I guess the model was good enough to  jump up in the private LB. Unfortunately, my XGBoost was still running at the cut off time and wasn't able to my blends unto date. \nSince the beginning I was using the full dataset as I did have access to a better machine with lots of RAM and swap space. Almost until last week, I was experimenting with the following models:</p>\n\n<ol>\n<li><p><a href=\"https://www.kaggle.com/scirpus\">@Scirpus</a> favourite FFM: I spent a lot of time studying the Avazu solution was trying to replicate it. In addition I also stumbled on the DeepFM paper (<a href=\"https://arxiv.org/abs/1703.04247\">https://arxiv.org/abs/1703.04247</a>) they have the code on GitHub (<a href=\"https://github.com/ChenglongChen/tensorflow-DeepFM\">https://github.com/ChenglongChen/tensorflow-DeepFM</a>). I wasn't able to tune the model beyond 0.96-0.97. I think a lot of my experimentation time was wasted chasing it and I don't think I am any closer to building better intuition around it even now. I will continue to spend time to study more. </p></li>\n<li><p>Vowpal Wabbit: This was another time sink in getting it to work. This was the most frustrating part, absolutely none of my models went beyond 0.47-0.48. I tried doing bayesian optimisation on it as well, but something was missing....maybe I did something wrong. </p></li>\n<li><p>Neural Network: I started of initially with the approach given by <a href=\"https://www.kaggle.com/alexanderkireev\">@Alexander Kireev</a>. I also did a bayesian optimisation to tune the layers and other NN parameters. These models got me in 0.979 range and adding more features from various kernels (next click, previous click and the subsequent combinations) started to bump me up the LB, but not by much. </p></li>\n<li><p>GBT models (XGBoost, CatBoost, LightGBM): These were relatively faster to train, so last week was focusing on getting them to tune (bayesian optimisation) and start bumping up the list. Don't use cat boost on large models. It was either that the pool or the fit method that blows up in the memory and crashes without it reaching any conclusion. But XGboost and LightGBM remained strong till the end. I was able to do more submissions in the last week that I did before that. my best scoring did not do better in LB so that was not used for scoring, but I used that as one, then I would have gotten into 0.9821211 on the private. Thats a big lesson learned...dont make hasty judgements. </p></li>\n</ol>\n\n<p>With 3 and 4 and the public word batch kernel was my final ensemble. I will share my notebook in the upcoming days. </p>\n\n<p>I did explore <a href=\"https://www.featuretools.com/\">feature tools</a> initially, but then relied more on the starting features in various kernels and then making more on top of them. </p>\n\n<p>Special Thanks to <a href=\"https://www.kaggle.com/yuliagm\">@yulia</a> <a href=\"https://www.kaggle.com/nanomathias\">@Nanomathias</a> <a href=\"https://www.kaggle.com/asraful70\">@Md Asraful Kabir</a> @Wenjie Bai <a href=\"https://www.kaggle.com/ogrellier\"></a><a href=\"/olivier\">@olivier</a> and CPMP</p>\n\n<p>Congratulations to the prize winners and all those who worked hard towards this competition. </p>\n\n<p>The learning journey continues. Note to myself, Make better notes, experiment more. </p>",
  "messages": [
    {
      "id": 326215,
      "postDate": "2018-05-09T12:58:56.140Z",
      "content": "<p>This was a crazy competition with the release of last minute of 9811 kernel, I stayed up the night to retune my models and see if I can beat that score. I edged it out only slightly, but I guess the model was good enough to  jump up in the private LB. Unfortunately, my XGBoost was still running at the cut off time and wasn't able to my blends unto date. \nSince the beginning I was using the full dataset as I did have access to a better machine with lots of RAM and swap space. Almost until last week, I was experimenting with the following models:</p>\n\n<ol>\n<li><p><a href=\"https://www.kaggle.com/scirpus\">@Scirpus</a> favourite FFM: I spent a lot of time studying the Avazu solution was trying to replicate it. In addition I also stumbled on the DeepFM paper (<a href=\"https://arxiv.org/abs/1703.04247\">https://arxiv.org/abs/1703.04247</a>) they have the code on GitHub (<a href=\"https://github.com/ChenglongChen/tensorflow-DeepFM\">https://github.com/ChenglongChen/tensorflow-DeepFM</a>). I wasn't able to tune the model beyond 0.96-0.97. I think a lot of my experimentation time was wasted chasing it and I don't think I am any closer to building better intuition around it even now. I will continue to spend time to study more. </p></li>\n<li><p>Vowpal Wabbit: This was another time sink in getting it to work. This was the most frustrating part, absolutely none of my models went beyond 0.47-0.48. I tried doing bayesian optimisation on it as well, but something was missing....maybe I did something wrong. </p></li>\n<li><p>Neural Network: I started of initially with the approach given by <a href=\"https://www.kaggle.com/alexanderkireev\">@Alexander Kireev</a>. I also did a bayesian optimisation to tune the layers and other NN parameters. These models got me in 0.979 range and adding more features from various kernels (next click, previous click and the subsequent combinations) started to bump me up the LB, but not by much. </p></li>\n<li><p>GBT models (XGBoost, CatBoost, LightGBM): These were relatively faster to train, so last week was focusing on getting them to tune (bayesian optimisation) and start bumping up the list. Don't use cat boost on large models. It was either that the pool or the fit method that blows up in the memory and crashes without it reaching any conclusion. But XGboost and LightGBM remained strong till the end. I was able to do more submissions in the last week that I did before that. my best scoring did not do better in LB so that was not used for scoring, but I used that as one, then I would have gotten into 0.9821211 on the private. Thats a big lesson learned...dont make hasty judgements. </p></li>\n</ol>\n\n<p>With 3 and 4 and the public word batch kernel was my final ensemble. I will share my notebook in the upcoming days. </p>\n\n<p>I did explore <a href=\"https://www.featuretools.com/\">feature tools</a> initially, but then relied more on the starting features in various kernels and then making more on top of them. </p>\n\n<p>Special Thanks to <a href=\"https://www.kaggle.com/yuliagm\">@yulia</a> <a href=\"https://www.kaggle.com/nanomathias\">@Nanomathias</a> <a href=\"https://www.kaggle.com/asraful70\">@Md Asraful Kabir</a> @Wenjie Bai <a href=\"https://www.kaggle.com/ogrellier\"></a><a href=\"/olivier\">@olivier</a> and CPMP</p>\n\n<p>Congratulations to the prize winners and all those who worked hard towards this competition. </p>\n\n<p>The learning journey continues. Note to myself, Make better notes, experiment more. </p>",
      "rawMarkdown": "This was a crazy competition with the release of last minute of 9811 kernel, I stayed up the night to retune my models and see if I can beat that score. I edged it out only slightly, but I guess the model was good enough to  jump up in the private LB. Unfortunately, my XGBoost was still running at the cut off time and wasn't able to my blends unto date. \nSince the beginning I was using the full dataset as I did have access to a better machine with lots of RAM and swap space. Almost until last week, I was experimenting with the following models:\n\n1. [@Scirpus][1] favourite FFM: I spent a lot of time studying the Avazu solution was trying to replicate it. In addition I also stumbled on the DeepFM paper (https://arxiv.org/abs/1703.04247) they have the code on GitHub (https://github.com/ChenglongChen/tensorflow-DeepFM). I wasn't able to tune the model beyond 0.96-0.97. I think a lot of my experimentation time was wasted chasing it and I don't think I am any closer to building better intuition around it even now. I will continue to spend time to study more. \n\n2.  Vowpal Wabbit: This was another time sink in getting it to work. This was the most frustrating part, absolutely none of my models went beyond 0.47-0.48. I tried doing bayesian optimisation on it as well, but something was missing....maybe I did something wrong. \n\n3. Neural Network: I started of initially with the approach given by [@Alexander Kireev][2]. I also did a bayesian optimisation to tune the layers and other NN parameters. These models got me in 0.979 range and adding more features from various kernels (next click, previous click and the subsequent combinations) started to bump me up the LB, but not by much. \n\n4. GBT models (XGBoost, CatBoost, LightGBM): These were relatively faster to train, so last week was focusing on getting them to tune (bayesian optimisation) and start bumping up the list. Don't use cat boost on large models. It was either that the pool or the fit method that blows up in the memory and crashes without it reaching any conclusion. But XGboost and LightGBM remained strong till the end. I was able to do more submissions in the last week that I did before that. my best scoring did not do better in LB so that was not used for scoring, but I used that as one, then I would have gotten into 0.9821211 on the private. Thats a big lesson learned...dont make hasty judgements. \n\nWith 3 and 4 and the public word batch kernel was my final ensemble. I will share my notebook in the upcoming days. \n\nI did explore [feature tools][3] initially, but then relied more on the starting features in various kernels and then making more on top of them. \n\nSpecial Thanks to [@yulia][4] [@Nanomathias][5] [@Md Asraful Kabir][6] @Wenjie Bai [@olivier][7] and CPMP\n\nCongratulations to the prize winners and all those who worked hard towards this competition. \n\nThe learning journey continues. Note to myself, Make better notes, experiment more. \n\n\n  [1]: https://www.kaggle.com/scirpus\n  [2]: https://www.kaggle.com/alexanderkireev\n  [3]: https://www.featuretools.com\n  [4]: https://www.kaggle.com/yuliagm\n  [5]: https://www.kaggle.com/nanomathias\n  [6]: https://www.kaggle.com/asraful70\n  [7]: https://www.kaggle.com/ogrellier",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "326215": "This was a crazy competition with the release of last minute of 9811 kernel, I stayed up the night to retune my models and see if I can beat that score. I edged it out only slightly, but I guess the model was good enough to  jump up in the private LB. Unfortunately, my XGBoost was still running at the cut off time and wasn't able to my blends unto date. \nSince the beginning I was using the full dataset as I did have access to a better machine with lots of RAM and swap space. Almost until last week, I was experimenting with the following models:\n\n1. [@Scirpus][1] favourite FFM: I spent a lot of time studying the Avazu solution was trying to replicate it. In addition I also stumbled on the DeepFM paper (https://arxiv.org/abs/1703.04247) they have the code on GitHub (https://github.com/ChenglongChen/tensorflow-DeepFM). I wasn't able to tune the model beyond 0.96-0.97. I think a lot of my experimentation time was wasted chasing it and I don't think I am any closer to building better intuition around it even now. I will continue to spend time to study more. \n\n2.  Vowpal Wabbit: This was another time sink in getting it to work. This was the most frustrating part, absolutely none of my models went beyond 0.47-0.48. I tried doing bayesian optimisation on it as well, but something was missing....maybe I did something wrong. \n\n3. Neural Network: I started of initially with the approach given by [@Alexander Kireev][2]. I also did a bayesian optimisation to tune the layers and other NN parameters. These models got me in 0.979 range and adding more features from various kernels (next click, previous click and the subsequent combinations) started to bump me up the LB, but not by much. \n\n4. GBT models (XGBoost, CatBoost, LightGBM): These were relatively faster to train, so last week was focusing on getting them to tune (bayesian optimisation) and start bumping up the list. Don't use cat boost on large models. It was either that the pool or the fit method that blows up in the memory and crashes without it reaching any conclusion. But XGboost and LightGBM remained strong till the end. I was able to do more submissions in the last week that I did before that. my best scoring did not do better in LB so that was not used for scoring, but I used that as one, then I would have gotten into 0.9821211 on the private. Thats a big lesson learned...dont make hasty judgements. \n\nWith 3 and 4 and the public word batch kernel was my final ensemble. I will share my notebook in the upcoming days. \n\nI did explore [feature tools][3] initially, but then relied more on the starting features in various kernels and then making more on top of them. \n\nSpecial Thanks to [@yulia][4] [@Nanomathias][5] [@Md Asraful Kabir][6] @Wenjie Bai [@olivier][7] and CPMP\n\nCongratulations to the prize winners and all those who worked hard towards this competition. \n\nThe learning journey continues. Note to myself, Make better notes, experiment more. \n\n\n  [1]: https://www.kaggle.com/scirpus\n  [2]: https://www.kaggle.com/alexanderkireev\n  [3]: https://www.featuretools.com\n  [4]: https://www.kaggle.com/yuliagm\n  [5]: https://www.kaggle.com/nanomathias\n  [6]: https://www.kaggle.com/asraful70\n  [7]: https://www.kaggle.com/ogrellier"
  }
}