{
  "id": 25521,
  "title": "Interesting solutions from the past.!",
  "url": "/competitions/outbrain-click-prediction/discussion/25521",
  "author_name": "",
  "post_date": "2016-11-16T18:39:02.573Z",
  "votes": 35,
  "comment_count": 14,
  "views": 2687,
  "content": "<p>I was going through the solutions of the past click prediction competitions in Kaggle. Here are some interesting solution links from them. Hope this will help others as well. </p>\n\n<ol>\n<li><p><a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums\">Criteo Display Advertising Challenge</a></p>\n\n<p>a. 3 idiots <a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10555/3-idiots-solution-libffm\">solution using LibFFM</a></p>\n\n<p>b. Third place <a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10547/document-and-code-for-the-3rd-place-finish\">solution</a></p>\n\n<p>c. <a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory\">FTRL starter</a> by tinrtgu</p>\n\n<p>d. <a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/9583/beat-the-benchmark-with-vowpal-wabbit\">Vowpal Wabbit starter</a> by Triskelion</p></li>\n<li><p><a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums\">Avazu Click Through Rate Prediction</a></p>\n\n<p>a. 4 idiots <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/12608/4-idiots-solution-libffm\">solution</a></p>\n\n<p>b. Learnings from <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/12476/vowpal-wabbit-lessons\">vowpal wabbit</a></p>\n\n<p>c. Improved <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/10927/beat-the-benchmark-with-less-than-1mb-of-memory\">FTRL model</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums\">Avito Context Ad Clicks</a></p>\n\n<p>a. First place <a href=\"https://github.com/owenzhang/kaggle-avito\">solution</a> By Owen</p>\n\n<p>b. Second place <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/16084/2st-place-winner-solution-gzs-iceberg\">solution</a></p>\n\n<p>c. Third place <a href=\"https://github.com/diefimov/avito_context_click_2015\">solution</a></p>\n\n<p>d. Solution sharing forum <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations\">post</a></p></li>\n</ol>",
  "messages": [
    {
      "id": "145095",
      "postDate": "11/16/2016 18:39:02",
      "content": "<p>I was going through the solutions of the past click prediction competitions in Kaggle. Here are some interesting solution links from them. Hope this will help others as well. </p>\n\n<ol>\n<li><p><a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums\">Criteo Display Advertising Challenge</a></p>\n\n<p>a. 3 idiots <a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10555/3-idiots-solution-libffm\">solution using LibFFM</a></p>\n\n<p>b. Third place <a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10547/document-and-code-for-the-3rd-place-finish\">solution</a></p>\n\n<p>c. <a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory\">FTRL starter</a> by tinrtgu</p>\n\n<p>d. <a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/9583/beat-the-benchmark-with-vowpal-wabbit\">Vowpal Wabbit starter</a> by Triskelion</p></li>\n<li><p><a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums\">Avazu Click Through Rate Prediction</a></p>\n\n<p>a. 4 idiots <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/12608/4-idiots-solution-libffm\">solution</a></p>\n\n<p>b. Learnings from <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/12476/vowpal-wabbit-lessons\">vowpal wabbit</a></p>\n\n<p>c. Improved <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/10927/beat-the-benchmark-with-less-than-1mb-of-memory\">FTRL model</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums\">Avito Context Ad Clicks</a></p>\n\n<p>a. First place <a href=\"https://github.com/owenzhang/kaggle-avito\">solution</a> By Owen</p>\n\n<p>b. Second place <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/16084/2st-place-winner-solution-gzs-iceberg\">solution</a></p>\n\n<p>c. Third place <a href=\"https://github.com/diefimov/avito_context_click_2015\">solution</a></p>\n\n<p>d. Solution sharing forum <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations\">post</a></p></li>\n</ol>",
      "rawMarkdown": "I was going through the solutions of the past click prediction competitions in Kaggle. Here are some interesting solution links from them. Hope this will help others as well. \r\n\r\n1. [Criteo Display Advertising Challenge][1]\r\n\r\n    a. 3 idiots [solution using LibFFM][2]\r\n\r\n    b. Third place [solution][3]\r\n    \r\n    c. [FTRL starter][4] by tinrtgu\r\n\r\n    d. [Vowpal Wabbit starter][5] by Triskelion\r\n\r\n2. [Avazu Click Through Rate Prediction][6]\r\n\r\n    a. 4 idiots [solution][7]\r\n  \r\n    b. Learnings from [vowpal wabbit][8]\r\n\r\n    c. Improved [FTRL model][9]\r\n\r\n3. [Avito Context Ad Clicks][10]\r\n\r\n    a. First place [solution][11] By Owen\r\n\r\n    b. Second place [solution][12]\r\n\r\n    c. Third place [solution][13]\r\n\r\n    d. Solution sharing forum [post][14]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums\r\n  [2]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10555/3-idiots-solution-libffm\r\n  [3]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10547/document-and-code-for-the-3rd-place-finish\r\n  [4]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory\r\n  [5]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/9583/beat-the-benchmark-with-vowpal-wabbit\r\n  [6]: https://www.kaggle.com/c/avazu-ctr-prediction/forums\r\n  [7]: https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/12608/4-idiots-solution-libffm\r\n  [8]: https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/12476/vowpal-wabbit-lessons\r\n  [9]: https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/10927/beat-the-benchmark-with-less-than-1mb-of-memory\r\n  [10]: https://www.kaggle.com/c/avito-context-ad-clicks/forums\r\n  [11]: https://github.com/owenzhang/kaggle-avito\r\n  [12]: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/16084/2st-place-winner-solution-gzs-iceberg\r\n  [13]: https://github.com/diefimov/avito_context_click_2015\r\n  [14]: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations",
      "votes": null
    },
    {
      "id": "145281",
      "postDate": "11/17/2016 17:04:37",
      "content": "<p>@SRK, thank you very much! I am most interested in the context ad one. is it any similar to this contest?</p>",
      "rawMarkdown": "SRK, thank you very much! I am most interested in the context ad one. is it any similar to this contest?",
      "votes": null
    },
    {
      "id": "145287",
      "postDate": "11/17/2016 17:31:14",
      "content": "<p>@Carl : Yes, Avito context ad is quite similar. Though the eval metric for that one is log loss (we can also convert this problem from recommendation to binary one as you know), the nature of the problem and the dataset are highly similar. I think we can get many insights from that particular competition's solution.</p>\n\n<p>If my memory serves me right, I got an opportunity to team up with one of the current leaders in that competition ;) </p>",
      "rawMarkdown": "Carl : Yes, Avito context ad is quite similar. Though the eval metric for that one is log loss (we can also convert this problem from recommendation to binary one as you know), the nature of the problem and the dataset are highly similar. I think we can get many insights from that particular competition's solution.\r\n\r\nIf my memory serves me right, I got an opportunity to team up with one of the current leaders in that competition ;)",
      "votes": null
    },
    {
      "id": "145293",
      "postDate": "11/17/2016 17:49:19",
      "content": "<p>@SRK: thanks for aggregating the past ones. I've taken a trip down memory lane for Criteo and Avazu, but the Avito didn't occur to (as i have not taken part in that one). Main message (from the first two at least): feature engineering + FFM is the killer combo. </p>\n\n<p>Speaking of FFM: can anybody suggest a good way of converting tabular data (with categorical columns) into libsvm format? The only one i've figured out so far is by converting to a sparse matrix first, but it's not exactly very efficient.</p>",
      "rawMarkdown": "SRK: thanks for aggregating the past ones. I've taken a trip down memory lane for Criteo and Avazu, but the Avito didn't occur to (as i have not taken part in that one). Main message (from the first two at least): feature engineering + FFM is the killer combo. \r\n\r\nSpeaking of FFM: can anybody suggest a good way of converting tabular data (with categorical columns) into libsvm format? The only one i've figured out so far is by converting to a sparse matrix first, but it's not exactly very efficient.",
      "votes": null
    },
    {
      "id": "145298",
      "postDate": "11/17/2016 18:10:51",
      "content": "<p>@Konrad : I am not sure I got you right. Can we hash the categorical columns instead of converting it to a sparse matrix? </p>",
      "rawMarkdown": "Konrad : I am not sure I got you right. Can we hash the categorical columns instead of converting it to a sparse matrix?",
      "votes": null
    },
    {
      "id": "145301",
      "postDate": "11/17/2016 18:12:05",
      "content": "<p>(facepalm): hash first, then create the features for FFM, of course ... Thanks you @SRK - i know i was missing something important in the scheme of things :-)</p>",
      "rawMarkdown": "(facepalm): hash first, then create the features for FFM, of course ... Thanks you @SRK - i know i was missing something important in the scheme of things :-)",
      "votes": null
    },
    {
      "id": "145304",
      "postDate": "11/17/2016 18:19:17",
      "content": "<p>Welcome @Konrad.. It happens at times :-) </p>",
      "rawMarkdown": "Welcome @Konrad.. It happens at times :-)",
      "votes": null
    },
    {
      "id": "145313",
      "postDate": "11/17/2016 19:28:11",
      "content": "<p>@SRK, how can I forget that we were in the same team @ avito?! :P</p>",
      "rawMarkdown": "SRK, how can I forget that we were in the same team @ avito?! :P",
      "votes": null
    },
    {
      "id": "145466",
      "postDate": "11/18/2016 14:21:08",
      "content": "<p>But really, you can write simple python script which do that even without using hashing trick. To do that with 1,2mln features I needed ~200MB RAM, and to do that with 16mln features I needed ~2GB RAM. I used python dictionary to store the key-value and pypy to speed up the computation. I attached some simple version. </p>",
      "rawMarkdown": "But really, you can write simple python script which do that even without using hashing trick. To do that with 1,2mln features I needed ~200MB RAM, and to do that with 16mln features I needed ~2GB RAM. I used python dictionary to store the key-value and pypy to speed up the computation. I attached some simple version.",
      "votes": null
    },
    {
      "id": "145720",
      "postDate": "11/20/2016 15:49:08",
      "content": "<p>Is the libffm format the same as libsvm? it seems like it's:</p>\n\n<p><code>0/1 target col_num:row_num:value</code> ?</p>\n\n<p>libsvm is: <code>0/1 target  col_num:value</code> no?</p>\n\n<p>but slide 13 here:\n<a href=\"https://www.csie.ntu.edu.tw/~r01922136/slides/ffm.pdf\">https://www.csie.ntu.edu.tw/~r01922136/slides/ffm.pdf</a></p>\n\n<p>suggests that it is something totally different?</p>\n\n<p><code>col_num:weird_num:value</code></p>\n\n<p>where <code>weird_num</code> seems like it is a unique value to identify one hot encoded features.</p>",
      "rawMarkdown": "Is the libffm format the same as libsvm? it seems like it's:\r\n\r\n`0/1 target col_num:row_num:value` ?\r\n\r\nlibsvm is: `0/1 target  col_num:value` no?\r\n\r\nbut slide 13 here:\r\nhttps://www.csie.ntu.edu.tw/~r01922136/slides/ffm.pdf\r\n\r\nsuggests that it is something totally different?\r\n\r\n`col_num:weird_num:value`\r\n\r\nwhere `weird_num` seems like it is a unique value to identify one hot encoded features.",
      "votes": null
    },
    {
      "id": "148445",
      "postDate": "12/05/2016 02:35:11",
      "content": "<p>Guess XGBoost and Libffm are the two models that I should try? </p>",
      "rawMarkdown": "Guess XGBoost and Libffm are the two models that I should try?",
      "votes": null
    },
    {
      "id": "154003",
      "postDate": "01/04/2017 11:42:58",
      "content": "<p>Has anyone being successful incorporating GBDT Leaf predictions as features for FFM, as in 3 idiots approach for Criteo competition?\nI tried many XGBoost and LightGBM configurations (e.g. 30 trees with depth 6 and 64 leafs each) but my results on FFM are worse both on training and validation sets.</p>",
      "rawMarkdown": "Has anyone being successful incorporating GBDT Leaf predictions as features for FFM, as in 3 idiots approach for Criteo competition?\r\nI tried many XGBoost and LightGBM configurations (e.g. 30 trees with depth 6 and 64 leafs each) but my results on FFM are worse both on training and validation sets.",
      "votes": null
    },
    {
      "id": "154005",
      "postDate": "01/04/2017 11:47:08",
      "content": "<p>@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical </p>",
      "rawMarkdown": "Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical",
      "votes": null
    },
    {
      "id": "154514",
      "postDate": "01/06/2017 15:43:14",
      "content": "<p>[quote=Keerath Jaggi;154005]</p>\n\n<p>@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical </p>\n\n<p>[/quote]</p>\n\n<p>what is the reason behind this explication? can you clarify please? Thanks!</p>",
      "rawMarkdown": "[quote=Keerath Jaggi;154005]\r\n\r\n@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical \r\n\r\n[/quote]\r\n\r\nwhat is the reason behind this explication? can you clarify please? Thanks!",
      "votes": null
    },
    {
      "id": "154581",
      "postDate": "01/06/2017 22:54:41",
      "content": "<p>[quote=Keerath Jaggi;154005]</p>\n\n<p>@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical </p>\n\n<p>[/quote]</p>\n\n<p>GBDT + numerical feature  -&gt; binning </p>\n\n<p>GBDT + categorical feature -&gt; feature transformation/reduction</p>\n\n<p>Maybe the former is more useful but latter is also ok.</p>",
      "rawMarkdown": "[quote=Keerath Jaggi;154005]\r\n\r\n@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical \r\n\r\n[/quote]\r\n\r\nGBDT + numerical feature  -> binning \r\n\r\nGBDT + categorical feature -> feature transformation/reduction\r\n\r\nMaybe the former is more useful but latter is also ok.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 145281,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/17/2016 17:04:37",
      "content": "<p>@SRK, thank you very much! I am most interested in the context ad one. is it any similar to this contest?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145287,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "11/17/2016 17:31:14",
      "content": "<p>@Carl : Yes, Avito context ad is quite similar. Though the eval metric for that one is log loss (we can also convert this problem from recommendation to binary one as you know), the nature of the problem and the dataset are highly similar. I think we can get many insights from that particular competition's solution.</p>\n\n<p>If my memory serves me right, I got an opportunity to team up with one of the current leaders in that competition ;) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145293,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "11/17/2016 17:49:19",
      "content": "<p>@SRK: thanks for aggregating the past ones. I've taken a trip down memory lane for Criteo and Avazu, but the Avito didn't occur to (as i have not taken part in that one). Main message (from the first two at least): feature engineering + FFM is the killer combo. </p>\n\n<p>Speaking of FFM: can anybody suggest a good way of converting tabular data (with categorical columns) into libsvm format? The only one i've figured out so far is by converting to a sparse matrix first, but it's not exactly very efficient.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145298,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "11/17/2016 18:10:51",
      "content": "<p>@Konrad : I am not sure I got you right. Can we hash the categorical columns instead of converting it to a sparse matrix? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145301,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "11/17/2016 18:12:05",
      "content": "<p>(facepalm): hash first, then create the features for FFM, of course ... Thanks you @SRK - i know i was missing something important in the scheme of things :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145304,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "11/17/2016 18:19:17",
      "content": "<p>Welcome @Konrad.. It happens at times :-) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145313,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "11/17/2016 19:28:11",
      "content": "<p>@SRK, how can I forget that we were in the same team @ avito?! :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145466,
      "author_name": "adamszalucha",
      "author_url": "",
      "post_date": "11/18/2016 14:21:08",
      "content": "<p>But really, you can write simple python script which do that even without using hashing trick. To do that with 1,2mln features I needed ~200MB RAM, and to do that with 16mln features I needed ~2GB RAM. I used python dictionary to store the key-value and pypy to speed up the computation. I attached some simple version. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 145720,
      "author_name": "brianlaw",
      "author_url": "",
      "post_date": "11/20/2016 15:49:08",
      "content": "<p>Is the libffm format the same as libsvm? it seems like it's:</p>\n\n<p><code>0/1 target col_num:row_num:value</code> ?</p>\n\n<p>libsvm is: <code>0/1 target  col_num:value</code> no?</p>\n\n<p>but slide 13 here:\n<a href=\"https://www.csie.ntu.edu.tw/~r01922136/slides/ffm.pdf\">https://www.csie.ntu.edu.tw/~r01922136/slides/ffm.pdf</a></p>\n\n<p>suggests that it is something totally different?</p>\n\n<p><code>col_num:weird_num:value</code></p>\n\n<p>where <code>weird_num</code> seems like it is a unique value to identify one hot encoded features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 148445,
      "author_name": "peixiang",
      "author_url": "",
      "post_date": "12/05/2016 02:35:11",
      "content": "<p>Guess XGBoost and Libffm are the two models that I should try? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154003,
      "author_name": "gspmoreira",
      "author_url": "",
      "post_date": "01/04/2017 11:42:58",
      "content": "<p>Has anyone being successful incorporating GBDT Leaf predictions as features for FFM, as in 3 idiots approach for Criteo competition?\nI tried many XGBoost and LightGBM configurations (e.g. 30 trees with depth 6 and 64 leafs each) but my results on FFM are worse both on training and validation sets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 154005,
      "author_name": "keerath",
      "author_url": "",
      "post_date": "01/04/2017 11:47:08",
      "content": "<p>@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical </p>",
      "votes": null,
      "replies": [
        {
          "id": 154514,
          "author_name": "betterplace",
          "author_url": "",
          "post_date": "01/06/2017 15:43:14",
          "content": "<p>[quote=Keerath Jaggi;154005]</p>\n\n<p>@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical </p>\n\n<p>[/quote]</p>\n\n<p>what is the reason behind this explication? can you clarify please? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 154581,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "01/06/2017 22:54:41",
          "content": "<p>[quote=Keerath Jaggi;154005]</p>\n\n<p>@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical </p>\n\n<p>[/quote]</p>\n\n<p>GBDT + numerical feature  -&gt; binning </p>\n\n<p>GBDT + categorical feature -&gt; feature transformation/reduction</p>\n\n<p>Maybe the former is more useful but latter is also ok.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "145095": "I was going through the solutions of the past click prediction competitions in Kaggle. Here are some interesting solution links from them. Hope this will help others as well. \r\n\r\n1. [Criteo Display Advertising Challenge][1]\r\n\r\n    a. 3 idiots [solution using LibFFM][2]\r\n\r\n    b. Third place [solution][3]\r\n    \r\n    c. [FTRL starter][4] by tinrtgu\r\n\r\n    d. [Vowpal Wabbit starter][5] by Triskelion\r\n\r\n2. [Avazu Click Through Rate Prediction][6]\r\n\r\n    a. 4 idiots [solution][7]\r\n  \r\n    b. Learnings from [vowpal wabbit][8]\r\n\r\n    c. Improved [FTRL model][9]\r\n\r\n3. [Avito Context Ad Clicks][10]\r\n\r\n    a. First place [solution][11] By Owen\r\n\r\n    b. Second place [solution][12]\r\n\r\n    c. Third place [solution][13]\r\n\r\n    d. Solution sharing forum [post][14]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums\r\n  [2]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10555/3-idiots-solution-libffm\r\n  [3]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10547/document-and-code-for-the-3rd-place-finish\r\n  [4]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory\r\n  [5]: https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/9583/beat-the-benchmark-with-vowpal-wabbit\r\n  [6]: https://www.kaggle.com/c/avazu-ctr-prediction/forums\r\n  [7]: https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/12608/4-idiots-solution-libffm\r\n  [8]: https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/12476/vowpal-wabbit-lessons\r\n  [9]: https://www.kaggle.com/c/avazu-ctr-prediction/forums/t/10927/beat-the-benchmark-with-less-than-1mb-of-memory\r\n  [10]: https://www.kaggle.com/c/avito-context-ad-clicks/forums\r\n  [11]: https://github.com/owenzhang/kaggle-avito\r\n  [12]: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/16084/2st-place-winner-solution-gzs-iceberg\r\n  [13]: https://github.com/diefimov/avito_context_click_2015\r\n  [14]: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations",
    "145281": "SRK, thank you very much! I am most interested in the context ad one. is it any similar to this contest?",
    "145287": "Carl : Yes, Avito context ad is quite similar. Though the eval metric for that one is log loss (we can also convert this problem from recommendation to binary one as you know), the nature of the problem and the dataset are highly similar. I think we can get many insights from that particular competition's solution.\r\n\r\nIf my memory serves me right, I got an opportunity to team up with one of the current leaders in that competition ;)",
    "145293": "SRK: thanks for aggregating the past ones. I've taken a trip down memory lane for Criteo and Avazu, but the Avito didn't occur to (as i have not taken part in that one). Main message (from the first two at least): feature engineering + FFM is the killer combo. \r\n\r\nSpeaking of FFM: can anybody suggest a good way of converting tabular data (with categorical columns) into libsvm format? The only one i've figured out so far is by converting to a sparse matrix first, but it's not exactly very efficient.",
    "145298": "Konrad : I am not sure I got you right. Can we hash the categorical columns instead of converting it to a sparse matrix?",
    "145301": "(facepalm): hash first, then create the features for FFM, of course ... Thanks you @SRK - i know i was missing something important in the scheme of things :-)",
    "145304": "Welcome @Konrad.. It happens at times :-)",
    "145313": "SRK, how can I forget that we were in the same team @ avito?! :P",
    "145466": "But really, you can write simple python script which do that even without using hashing trick. To do that with 1,2mln features I needed ~200MB RAM, and to do that with 16mln features I needed ~2GB RAM. I used python dictionary to store the key-value and pypy to speed up the computation. I attached some simple version.",
    "145720": "Is the libffm format the same as libsvm? it seems like it's:\r\n\r\n`0/1 target col_num:row_num:value` ?\r\n\r\nlibsvm is: `0/1 target  col_num:value` no?\r\n\r\nbut slide 13 here:\r\nhttps://www.csie.ntu.edu.tw/~r01922136/slides/ffm.pdf\r\n\r\nsuggests that it is something totally different?\r\n\r\n`col_num:weird_num:value`\r\n\r\nwhere `weird_num` seems like it is a unique value to identify one hot encoded features.",
    "148445": "Guess XGBoost and Libffm are the two models that I should try?",
    "154003": "Has anyone being successful incorporating GBDT Leaf predictions as features for FFM, as in 3 idiots approach for Criteo competition?\r\nI tried many XGBoost and LightGBM configurations (e.g. 30 trees with depth 6 and 64 leafs each) but my results on FFM are worse both on training and validation sets.",
    "154005": "Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical",
    "154514": "[quote=Keerath Jaggi;154005]\r\n\r\n@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical \r\n\r\n[/quote]\r\n\r\nwhat is the reason behind this explication? can you clarify please? Thanks!",
    "154581": "[quote=Keerath Jaggi;154005]\r\n\r\n@Gabriel GBDT based leaf score pred is useful when we have numerical features as opposed to this case where bulk of the features are categorical \r\n\r\n[/quote]\r\n\r\nGBDT + numerical feature  -> binning \r\n\r\nGBDT + categorical feature -> feature transformation/reduction\r\n\r\nMaybe the former is more useful but latter is also ok."
  },
  "source": "meta"
}