{
  "id": 56282,
  "title": "Best LibFFM",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56282",
  "author_name": "",
  "post_date": "2018-05-08T08:29:05.321637200Z",
  "votes": 27,
  "comment_count": 8,
  "views": 0,
  "content": "<p>0.9818 (private) was our best libffm</p>\n\n<p>To do this we had to take into account the unbalanced dataset - the simplest solution was to modify the kappa in the train function</p>\n\n<pre><code>ffm_float kappa = -y*expnyt/(1+expnyt);\nif(y == 1)kappa*=9;\nwTx(begin, end, r, *model, kappa, param.eta, param.lambda, true);\n</code></pre>\n\n<p>This gave an uplift of .001</p>",
  "messages": [
    {
      "id": "325217",
      "postDate": "05/08/2018 08:29:05",
      "content": "<p>0.9818 (private) was our best libffm</p>\n\n<p>To do this we had to take into account the unbalanced dataset - the simplest solution was to modify the kappa in the train function</p>\n\n<pre><code>ffm_float kappa = -y*expnyt/(1+expnyt);\nif(y == 1)kappa*=9;\nwTx(begin, end, r, *model, kappa, param.eta, param.lambda, true);\n</code></pre>\n\n<p>This gave an uplift of .001</p>",
      "rawMarkdown": "0.9818 (private) was our best libffm\n\nTo do this we had to take into account the unbalanced dataset - the simplest solution was to modify the kappa in the train function\n\n    ffm_float kappa = -y*expnyt/(1+expnyt);\n    if(y == 1)kappa*=9;\n    wTx(begin, end, r, *model, kappa, param.eta, param.lambda, true);\n\nThis gave an uplift of .001",
      "votes": null
    },
    {
      "id": "325222",
      "postDate": "05/08/2018 08:34:42",
      "content": "<p>Well done, congrats on the result.</p>",
      "rawMarkdown": "Well done, congrats on the result.",
      "votes": null
    },
    {
      "id": "325293",
      "postDate": "05/08/2018 09:27:42",
      "content": "<p>Hi Scirpus, interesting to see that field-aware factorization machine (FFM) works so well. \nWell done and thanks for the insight</p>",
      "rawMarkdown": "Hi Scirpus, interesting to see that field-aware factorization machine (FFM) works so well. \nWell done and thanks for the insight",
      "votes": null
    },
    {
      "id": "325595",
      "postDate": "05/08/2018 15:55:43",
      "content": "<p>Great to see you share your trick here :D</p>",
      "rawMarkdown": "Great to see you share your trick here :D",
      "votes": null
    },
    {
      "id": "325608",
      "postDate": "05/08/2018 16:15:33",
      "content": "<p>Good result and thanks for sharing Scripus</p>",
      "rawMarkdown": "Good result and thanks for sharing Scripus",
      "votes": null
    },
    {
      "id": "325614",
      "postDate": "05/08/2018 16:22:37",
      "content": "<p>LOL - yeah I did partner - I cannot keep secrets for long ;)</p>",
      "rawMarkdown": "LOL - yeah I did partner - I cannot keep secrets for long ;)",
      "votes": null
    },
    {
      "id": "325720",
      "postDate": "05/08/2018 19:35:13",
      "content": "<p>I tried my luck with FFM (xLearn) aswell but my results were much worse. Don't know the actual score but my best single model ffm should be around 0.9740 - 0.9750 (private LB). I basically rebuilt 3 idiots approach for Display Advertising Challenge, carefully building and tuning a Gradient Boosted Decision Tree with 30 trees and using its leaf index predictions as categorical features for my FFM. My FFM barely scored better than the GBDT with 30 trees :D But at least it contributed by some amount when blending with my LightGBM model I built in parallel.</p>\n\n<p>Further feature engineering and tuning was quite time consuming given the amount of data and the time it takes to prepare tabular data for FFM. I gave up on FFM at the end and focused on LightGBM.</p>\n\n<p>Would you mind sharing more details of your model and what features you used?</p>",
      "rawMarkdown": "I tried my luck with FFM (xLearn) aswell but my results were much worse. Don't know the actual score but my best single model ffm should be around 0.9740 - 0.9750 (private LB). I basically rebuilt 3 idiots approach for Display Advertising Challenge, carefully building and tuning a Gradient Boosted Decision Tree with 30 trees and using its leaf index predictions as categorical features for my FFM. My FFM barely scored better than the GBDT with 30 trees :D But at least it contributed by some amount when blending with my LightGBM model I built in parallel.\n\nFurther feature engineering and tuning was quite time consuming given the amount of data and the time it takes to prepare tabular data for FFM. I gave up on FFM at the end and focused on LightGBM.\n\nWould you mind sharing more details of your model and what features you used?",
      "votes": null
    },
    {
      "id": "325757",
      "postDate": "05/08/2018 20:35:32",
      "content": "<p>I binned continuous variables into 50 bins each.  Then used the kappa trick.  I used my modified ffm to display auc score to stop overfitting </p>",
      "rawMarkdown": "I binned continuous variables into 50 bins each.  Then used the kappa trick.  I used my modified ffm to display auc score to stop overfitting",
      "votes": null
    },
    {
      "id": "325891",
      "postDate": "05/09/2018 02:48:45",
      "content": "<p>Cool trick and thanks for sharing!</p>\n\n<p>Would you plan to share your project code on Kaggle or Github?</p>",
      "rawMarkdown": "Cool trick and thanks for sharing!\n\nWould you plan to share your project code on Kaggle or Github?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325222,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/08/2018 08:34:42",
      "content": "<p>Well done, congrats on the result.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325293,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/08/2018 09:27:42",
      "content": "<p>Hi Scirpus, interesting to see that field-aware factorization machine (FFM) works so well. \nWell done and thanks for the insight</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325595,
      "author_name": "cizonb",
      "author_url": "",
      "post_date": "05/08/2018 15:55:43",
      "content": "<p>Great to see you share your trick here :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 325614,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "05/08/2018 16:22:37",
          "content": "<p>LOL - yeah I did partner - I cannot keep secrets for long ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325608,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "05/08/2018 16:15:33",
      "content": "<p>Good result and thanks for sharing Scripus</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325720,
      "author_name": "pepeeee",
      "author_url": "",
      "post_date": "05/08/2018 19:35:13",
      "content": "<p>I tried my luck with FFM (xLearn) aswell but my results were much worse. Don't know the actual score but my best single model ffm should be around 0.9740 - 0.9750 (private LB). I basically rebuilt 3 idiots approach for Display Advertising Challenge, carefully building and tuning a Gradient Boosted Decision Tree with 30 trees and using its leaf index predictions as categorical features for my FFM. My FFM barely scored better than the GBDT with 30 trees :D But at least it contributed by some amount when blending with my LightGBM model I built in parallel.</p>\n\n<p>Further feature engineering and tuning was quite time consuming given the amount of data and the time it takes to prepare tabular data for FFM. I gave up on FFM at the end and focused on LightGBM.</p>\n\n<p>Would you mind sharing more details of your model and what features you used?</p>",
      "votes": null,
      "replies": [
        {
          "id": 325757,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "05/08/2018 20:35:32",
          "content": "<p>I binned continuous variables into 50 bins each.  Then used the kappa trick.  I used my modified ffm to display auc score to stop overfitting </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325891,
      "author_name": "shawnyxiao",
      "author_url": "",
      "post_date": "05/09/2018 02:48:45",
      "content": "<p>Cool trick and thanks for sharing!</p>\n\n<p>Would you plan to share your project code on Kaggle or Github?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325217": "0.9818 (private) was our best libffm\n\nTo do this we had to take into account the unbalanced dataset - the simplest solution was to modify the kappa in the train function\n\n    ffm_float kappa = -y*expnyt/(1+expnyt);\n    if(y == 1)kappa*=9;\n    wTx(begin, end, r, *model, kappa, param.eta, param.lambda, true);\n\nThis gave an uplift of .001",
    "325222": "Well done, congrats on the result.",
    "325293": "Hi Scirpus, interesting to see that field-aware factorization machine (FFM) works so well. \nWell done and thanks for the insight",
    "325595": "Great to see you share your trick here :D",
    "325608": "Good result and thanks for sharing Scripus",
    "325614": "LOL - yeah I did partner - I cannot keep secrets for long ;)",
    "325720": "I tried my luck with FFM (xLearn) aswell but my results were much worse. Don't know the actual score but my best single model ffm should be around 0.9740 - 0.9750 (private LB). I basically rebuilt 3 idiots approach for Display Advertising Challenge, carefully building and tuning a Gradient Boosted Decision Tree with 30 trees and using its leaf index predictions as categorical features for my FFM. My FFM barely scored better than the GBDT with 30 trees :D But at least it contributed by some amount when blending with my LightGBM model I built in parallel.\n\nFurther feature engineering and tuning was quite time consuming given the amount of data and the time it takes to prepare tabular data for FFM. I gave up on FFM at the end and focused on LightGBM.\n\nWould you mind sharing more details of your model and what features you used?",
    "325757": "I binned continuous variables into 50 bins each.  Then used the kappa trick.  I used my modified ffm to display auc score to stop overfitting",
    "325891": "Cool trick and thanks for sharing!\n\nWould you plan to share your project code on Kaggle or Github?"
  },
  "source": "meta"
}