{
  "id": 54620,
  "title": "Lack of model diversity in this competition?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54620",
  "author_name": "",
  "post_date": "2018-04-16T01:01:56.149023200Z",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>There seems to be a lack of diversity in discussion and kernels when it comes to modeling for this competitions. Everybody seems to be focused on feature engineering and uses LightGBM (some sporadic XGBoost). Only a few maestros shared their work with Wordbatch, which is enlightening. I am very surprised that nobody is talking about FFM, which is the king of CTR competitions in the past. </p>\n\n<p>I am new to Kaggle and this is my first CTR competitions. Is this the norm for CTR competitions?</p>",
  "messages": [
    {
      "id": "314628",
      "postDate": "04/16/2018 01:01:56",
      "content": "<p>There seems to be a lack of diversity in discussion and kernels when it comes to modeling for this competitions. Everybody seems to be focused on feature engineering and uses LightGBM (some sporadic XGBoost). Only a few maestros shared their work with Wordbatch, which is enlightening. I am very surprised that nobody is talking about FFM, which is the king of CTR competitions in the past. </p>\n\n<p>I am new to Kaggle and this is my first CTR competitions. Is this the norm for CTR competitions?</p>",
      "rawMarkdown": "There seems to be a lack of diversity in discussion and kernels when it comes to modeling for this competitions. Everybody seems to be focused on feature engineering and uses LightGBM (some sporadic XGBoost). Only a few maestros shared their work with Wordbatch, which is enlightening. I am very surprised that nobody is talking about FFM, which is the king of CTR competitions in the past. \n\nI am new to Kaggle and this is my first CTR competitions. Is this the norm for CTR competitions?",
      "votes": null
    },
    {
      "id": "314672",
      "postDate": "04/16/2018 04:38:54",
      "content": "<p>Yeah, I'm also new to kaggle and I heard FFM is powerful for CTR competition, check winner solution of a former CTR competition: <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/discussion/12608\">https://www.kaggle.com/c/avazu-ctr-prediction/discussion/12608</a>. And I think I will also try out <a href=\"https://github.com/aksnzhy/xlearn\">xlearn</a>.</p>",
      "rawMarkdown": "Yeah, I'm also new to kaggle and I heard FFM is powerful for CTR competition, check winner solution of a former CTR competition: https://www.kaggle.com/c/avazu-ctr-prediction/discussion/12608. And I think I will also try out [xlearn](https://github.com/aksnzhy/xlearn).",
      "votes": null
    },
    {
      "id": "315059",
      "postDate": "04/16/2018 17:47:15",
      "content": "<p>I think ffm is good at extracting embedding, but in this competition, it seems hard to generate proper ip feature\nhere is a kernel prepare libffm data but not so attractive than ooxx(0.97up)\n<a href=\"https://www.kaggle.com/mpearmain/pandas-to-libffm\">https://www.kaggle.com/mpearmain/pandas-to-libffm</a></p>",
      "rawMarkdown": "I think ffm is good at extracting embedding, but in this competition, it seems hard to generate proper ip feature\nhere is a kernel prepare libffm data but not so attractive than ooxx(0.97up)\nhttps://www.kaggle.com/mpearmain/pandas-to-libffm",
      "votes": null
    },
    {
      "id": "316224",
      "postDate": "04/18/2018 13:24:31",
      "content": "<p>Last time I used FFM (during the Outbrain challenge), training step was very RAM intensive. It was time consuming to test new features as well. I can't remember the exact dataset size but I think it was much smaller than this one.  </p>\n\n<p>I've not tried it yet because I am not sure my computer could handle it :) </p>",
      "rawMarkdown": "Last time I used FFM (during the Outbrain challenge), training step was very RAM intensive. It was time consuming to test new features as well. I can't remember the exact dataset size but I think it was much smaller than this one.  \n\nI've not tried it yet because I am not sure my computer could handle it :)",
      "votes": null
    },
    {
      "id": "316250",
      "postDate": "04/18/2018 14:35:18",
      "content": "<p>The trouble with FFM is that it's quite memory intensive on its own - and with the data size here, the problem is amplified. I think that's the main reason people are going for the online training of a regular FM (which is effectively what the Wordbatch kernels are doing).</p>",
      "rawMarkdown": "The trouble with FFM is that it's quite memory intensive on its own - and with the data size here, the problem is amplified. I think that's the main reason people are going for the online training of a regular FM (which is effectively what the Wordbatch kernels are doing).",
      "votes": null
    },
    {
      "id": "316358",
      "postDate": "04/18/2018 20:19:49",
      "content": "<p>Just in case your running into this issue with RAM ... one mistake I had been making with FFM was making the Field values too high... way higher than they needed to be. @alno explained to me that FFM allocates space according to the highest field value where using the convention <code>&lt;field1&gt;:&lt;feature1&gt;:&lt;value1&gt;</code> ... if you have 20 fields, make sure your highest field is id 20... not id 50 or something. Seems obvious but I made this mistake and it blew up my RAM. FTRL / Wordbatch seems to work similar -- if the hash value is 2 ** 30, it will create space for 2 ** 30 parameters, even most of these are never used....</p>\n\n<p>I imagine its the same for the feature values in FFM in each field - if you have 1000 features, make sure the highest feature value does not go over 1000. </p>",
      "rawMarkdown": "Just in case your running into this issue with RAM ... one mistake I had been making with FFM was making the Field values too high... way higher than they needed to be. @alno explained to me that FFM allocates space according to the highest field value where using the convention `",
      "votes": null
    },
    {
      "id": "316375",
      "postDate": "04/18/2018 21:17:16",
      "content": "<p>Thanks Darragh for the advice! I might have made this mistake, that gives me a good reason to try FFM again then.</p>",
      "rawMarkdown": "Thanks Darragh for the advice! I might have made this mistake, that gives me a good reason to try FFM again then.",
      "votes": null
    },
    {
      "id": "316706",
      "postDate": "04/19/2018 17:49:59",
      "content": "<p>I did not make any submissions yet, multi layer neural network with embedding of ip, app, os, channel seems to works pretty well in local validation (split randomly not by time). I got around 0.97+ without extra features.</p>",
      "rawMarkdown": "I did not make any submissions yet, multi layer neural network with embedding of ip, app, os, channel seems to works pretty well in local validation (split randomly not by time). I got around 0.97+ without extra features.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 314672,
      "author_name": "bangdasun",
      "author_url": "",
      "post_date": "04/16/2018 04:38:54",
      "content": "<p>Yeah, I'm also new to kaggle and I heard FFM is powerful for CTR competition, check winner solution of a former CTR competition: <a href=\"https://www.kaggle.com/c/avazu-ctr-prediction/discussion/12608\">https://www.kaggle.com/c/avazu-ctr-prediction/discussion/12608</a>. And I think I will also try out <a href=\"https://github.com/aksnzhy/xlearn\">xlearn</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 315059,
      "author_name": "lianglu",
      "author_url": "",
      "post_date": "04/16/2018 17:47:15",
      "content": "<p>I think ffm is good at extracting embedding, but in this competition, it seems hard to generate proper ip feature\nhere is a kernel prepare libffm data but not so attractive than ooxx(0.97up)\n<a href=\"https://www.kaggle.com/mpearmain/pandas-to-libffm\">https://www.kaggle.com/mpearmain/pandas-to-libffm</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 316224,
      "author_name": "phansoks",
      "author_url": "",
      "post_date": "04/18/2018 13:24:31",
      "content": "<p>Last time I used FFM (during the Outbrain challenge), training step was very RAM intensive. It was time consuming to test new features as well. I can't remember the exact dataset size but I think it was much smaller than this one.  </p>\n\n<p>I've not tried it yet because I am not sure my computer could handle it :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 316358,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "04/18/2018 20:19:49",
          "content": "<p>Just in case your running into this issue with RAM ... one mistake I had been making with FFM was making the Field values too high... way higher than they needed to be. @alno explained to me that FFM allocates space according to the highest field value where using the convention <code>&lt;field1&gt;:&lt;feature1&gt;:&lt;value1&gt;</code> ... if you have 20 fields, make sure your highest field is id 20... not id 50 or something. Seems obvious but I made this mistake and it blew up my RAM. FTRL / Wordbatch seems to work similar -- if the hash value is 2 ** 30, it will create space for 2 ** 30 parameters, even most of these are never used....</p>\n\n<p>I imagine its the same for the feature values in FFM in each field - if you have 1000 features, make sure the highest feature value does not go over 1000. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316375,
          "author_name": "phansoks",
          "author_url": "",
          "post_date": "04/18/2018 21:17:16",
          "content": "<p>Thanks Darragh for the advice! I might have made this mistake, that gives me a good reason to try FFM again then.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 316250,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "04/18/2018 14:35:18",
      "content": "<p>The trouble with FFM is that it's quite memory intensive on its own - and with the data size here, the problem is amplified. I think that's the main reason people are going for the online training of a regular FM (which is effectively what the Wordbatch kernels are doing).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 316706,
      "author_name": "lifepreserver",
      "author_url": "",
      "post_date": "04/19/2018 17:49:59",
      "content": "<p>I did not make any submissions yet, multi layer neural network with embedding of ip, app, os, channel seems to works pretty well in local validation (split randomly not by time). I got around 0.97+ without extra features.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "314628": "There seems to be a lack of diversity in discussion and kernels when it comes to modeling for this competitions. Everybody seems to be focused on feature engineering and uses LightGBM (some sporadic XGBoost). Only a few maestros shared their work with Wordbatch, which is enlightening. I am very surprised that nobody is talking about FFM, which is the king of CTR competitions in the past. \n\nI am new to Kaggle and this is my first CTR competitions. Is this the norm for CTR competitions?",
    "314672": "Yeah, I'm also new to kaggle and I heard FFM is powerful for CTR competition, check winner solution of a former CTR competition: https://www.kaggle.com/c/avazu-ctr-prediction/discussion/12608. And I think I will also try out [xlearn](https://github.com/aksnzhy/xlearn).",
    "315059": "I think ffm is good at extracting embedding, but in this competition, it seems hard to generate proper ip feature\nhere is a kernel prepare libffm data but not so attractive than ooxx(0.97up)\nhttps://www.kaggle.com/mpearmain/pandas-to-libffm",
    "316224": "Last time I used FFM (during the Outbrain challenge), training step was very RAM intensive. It was time consuming to test new features as well. I can't remember the exact dataset size but I think it was much smaller than this one.  \n\nI've not tried it yet because I am not sure my computer could handle it :)",
    "316250": "The trouble with FFM is that it's quite memory intensive on its own - and with the data size here, the problem is amplified. I think that's the main reason people are going for the online training of a regular FM (which is effectively what the Wordbatch kernels are doing).",
    "316358": "Just in case your running into this issue with RAM ... one mistake I had been making with FFM was making the Field values too high... way higher than they needed to be. @alno explained to me that FFM allocates space according to the highest field value where using the convention `",
    "316375": "Thanks Darragh for the advice! I might have made this mistake, that gives me a good reason to try FFM again then.",
    "316706": "I did not make any submissions yet, multi layer neural network with embedding of ip, app, os, channel seems to works pretty well in local validation (split randomly not by time). I got around 0.97+ without extra features."
  },
  "source": "meta"
}