{
  "id": 54929,
  "title": "any success with non tree based models?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54929",
  "author_name": "",
  "post_date": "2018-04-19T20:40:17.966387400Z",
  "votes": 17,
  "comment_count": 18,
  "views": 0,
  "content": "<p>i wonder how far you can go with other model types besdie lgb/xgb. so far i have only tried lgb and it is performing really well and relatively fast for such a big data. </p>",
  "messages": [
    {
      "id": "316759",
      "postDate": "04/19/2018 20:40:17",
      "content": "<p>i wonder how far you can go with other model types besdie lgb/xgb. so far i have only tried lgb and it is performing really well and relatively fast for such a big data. </p>",
      "rawMarkdown": "i wonder how far you can go with other model types besdie lgb/xgb. so far i have only tried lgb and it is performing really well and relatively fast for such a big data.",
      "votes": null
    },
    {
      "id": "316872",
      "postDate": "04/20/2018 05:22:44",
      "content": "<p>NN can also perform well (.977-.978)</p>",
      "rawMarkdown": "NN can also perform well (.977-.978)",
      "votes": null
    },
    {
      "id": "316936",
      "postDate": "04/20/2018 07:57:39",
      "content": "<p>but does it blend ? I have also a nn with .979x, it adds almost 0 to the ensemble. lgb ist at .981x. I feel that it must be tuned to equal levels like lgb, but this seems impossible.</p>",
      "rawMarkdown": "but does it blend ? I have also a nn with .979x, it adds almost 0 to the ensemble. lgb ist at .981x. I feel that it must be tuned to equal levels like lgb, but this seems impossible.",
      "votes": null
    },
    {
      "id": "316953",
      "postDate": "04/20/2018 08:45:25",
      "content": "<p>you are right, it adds very tiny to the ensemble (~.0001) </p>",
      "rawMarkdown": "you are right, it adds very tiny to the ensemble (~.0001)",
      "votes": null
    },
    {
      "id": "316984",
      "postDate": "04/20/2018 10:47:58",
      "content": "<p>Thanks Michael. is lbm same as Field wise Factorization Machine ? </p>",
      "rawMarkdown": "Thanks Michael. is lbm same as Field wise Factorization Machine ?",
      "votes": null
    },
    {
      "id": "316986",
      "postDate": "04/20/2018 10:50:17",
      "content": "<p>I think it is lightgbm rather ;)</p>",
      "rawMarkdown": "I think it is lightgbm rather ;)",
      "votes": null
    },
    {
      "id": "316988",
      "postDate": "04/20/2018 11:05:20",
      "content": "<p>Thanks @Michael for the info. If your NN is around 0.979, I won't even try NNs :)</p>",
      "rawMarkdown": "Thanks @Michael for the info. If your NN is around 0.979, I won't even try NNs :)",
      "votes": null
    },
    {
      "id": "316993",
      "postDate": "04/20/2018 11:38:04",
      "content": "<p>lbm  -&gt;  I mean lgb (light gbm) sorry!</p>",
      "rawMarkdown": "lbm  -&gt;  I mean lgb (light gbm) sorry!",
      "votes": null
    },
    {
      "id": "317553",
      "postDate": "04/21/2018 20:50:52",
      "content": "<p>I was able to achieve 0.9110 using a simple Naive Bayes model. \n(Which is far from \"performing well\" though, but hey, it's not tree based!)\n<a href=\"https://www.kaggle.com/ianchute/naive-bayes-model-on-eight-basic-features/notebook\">https://www.kaggle.com/ianchute/naive-bayes-model-on-eight-basic-features/notebook</a></p>",
      "rawMarkdown": "I was able to achieve 0.9110 using a simple Naive Bayes model. \n(Which is far from \"performing well\" though, but hey, it's not tree based!)\nhttps://www.kaggle.com/ianchute/naive-bayes-model-on-eight-basic-features/notebook",
      "votes": null
    },
    {
      "id": "318396",
      "postDate": "04/23/2018 18:42:33",
      "content": "<p>Seems like the only other model that comes close to LightGBM is FM_FTRL. From our experience, it also blends well.</p>",
      "rawMarkdown": "Seems like the only other model that comes close to LightGBM is FM_FTRL. From our experience, it also blends well.",
      "votes": null
    },
    {
      "id": "318413",
      "postDate": "04/23/2018 19:22:29",
      "content": "<p>CuteChibiko mentions a .9814 neural network (of unknown architecture) <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\">in this thread</a>, so it would seem there's still hope for deep learning. </p>",
      "rawMarkdown": "CuteChibiko mentions a .9814 neural network (of unknown architecture) [in this thread][1], so it would seem there's still hope for deep learning. \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634",
      "votes": null
    },
    {
      "id": "318429",
      "postDate": "04/23/2018 20:16:05",
      "content": "<p>Thank you <a href=\"/sijunhe\">@sijunhe</a>! May I ask what kind of parameters could be tuned for better performance in FM_FTRL?</p>",
      "rawMarkdown": "Thank you @sijunhe! May I ask what kind of parameters could be tuned for better performance in FM_FTRL?",
      "votes": null
    },
    {
      "id": "318430",
      "postDate": "04/23/2018 20:26:17",
      "content": "<blockquote>\n  <p><strong>Michael Jahrer wrote</strong></p>\n  \n  <blockquote>\n    <p>but does it blend ? I have also a nn with .979x, it adds almost 0 to the ensemble. lgb ist at .981x. I feel that it must be tuned to equal levels like lgb, but this seems impossible.</p>\n  </blockquote>\n</blockquote>\n\n<p>Have you tried using <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53689\">Generative Adversarial Denoising Autoencoder</a>?</p>",
      "rawMarkdown": "&gt; **Michael Jahrer wrote**\n&gt; \n&gt;&gt; but does it blend ? I have also a nn with .979x, it adds almost 0 to the ensemble. lgb ist at .981x. I feel that it must be tuned to equal levels like lgb, but this seems impossible.\n\nHave you tried using [Generative Adversarial Denoising Autoencoder][1]?\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53689",
      "votes": null
    },
    {
      "id": "318452",
      "postDate": "04/23/2018 21:00:08",
      "content": "<p>haha, no, you ?</p>\n\n<p>but yea, I spend some days in re-writing DAE code to make this data transformation on-the-fly (otherwise mem would explode). very very slow here, 1 epoch through 242M samples with a 1100-2000-2000-1100 net is at 2700[s]. No way to effectively tune something. Also I feel this sparse data is not suitable for DAE. supervised nn afterwards scored also not good, doesnt help in blend. Probably settingd/params are in wrong range, I see no light here :D</p>",
      "rawMarkdown": "haha, no, you ?\n\nbut yea, I spend some days in re-writing DAE code to make this data transformation on-the-fly (otherwise mem would explode). very very slow here, 1 epoch through 242M samples with a 1100-2000-2000-1100 net is at 2700[s]. No way to effectively tune something. Also I feel this sparse data is not suitable for DAE. supervised nn afterwards scored also not good, doesnt help in blend. Probably settingd/params are in wrong range, I see no light here :D",
      "votes": null
    },
    {
      "id": "318454",
      "postDate": "04/23/2018 21:03:24",
      "content": "<p>I tried gans and autoencoders (separately :p) and they were between .9799 and .9802, not really worth the effort.</p>",
      "rawMarkdown": "I tried gans and autoencoders (separately :p) and they were between .9799 and .9802, not really worth the effort.",
      "votes": null
    },
    {
      "id": "319036",
      "postDate": "04/25/2018 04:44:41",
      "content": "<p>Thanks for the heads up, Joe.</p>\n\n<p>@Snorlax, our FM_FTRL is about the same as the public kernel (~0.977). I didn't mean that FM_FTRL is as good as LGB at 0.981x., but it seemed like the model in public kernels with comparable performance. </p>",
      "rawMarkdown": "Thanks for the heads up, Joe.\n\n@Snorlax, our FM_FTRL is about the same as the public kernel (~0.977). I didn't mean that FM_FTRL is as good as LGB at 0.981x., but it seemed like the model in public kernels with comparable performance.",
      "votes": null
    },
    {
      "id": "319204",
      "postDate": "04/25/2018 13:52:10",
      "content": "<p>your single model lgb is at .981x, wow that's great, seems I should work harder :)</p>",
      "rawMarkdown": "your single model lgb is at .981x, wow that's great, seems I should work harder :)",
      "votes": null
    },
    {
      "id": "319416",
      "postDate": "04/26/2018 02:19:56",
      "content": "<p>Blend of 3 model (1 LGBM(0.9815) model and 2 NN(0.9815) models) got LB: 09819</p>\n\n<p>It's just weighted average. It seems work without autoencoders, :)</p>",
      "rawMarkdown": "Blend of 3 model (1 LGBM(0.9815) model and 2 NN(0.9815) models) got LB: 09819\n\nIt's just weighted average. It seems work without autoencoders, :)",
      "votes": null
    },
    {
      "id": "319480",
      "postDate": "04/26/2018 06:48:04",
      "content": "<p>@CuteChibiko, you convinced me to start on NN, so far working with a single lgb, I will soon reach a wall I think given how people ahead of me are stalled.</p>",
      "rawMarkdown": "CuteChibiko, you convinced me to start on NN, so far working with a single lgb, I will soon reach a wall I think given how people ahead of me are stalled.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 316872,
      "author_name": "hoogleraar",
      "author_url": "",
      "post_date": "04/20/2018 05:22:44",
      "content": "<p>NN can also perform well (.977-.978)</p>",
      "votes": null,
      "replies": [
        {
          "id": 316936,
          "author_name": "mjahrer",
          "author_url": "",
          "post_date": "04/20/2018 07:57:39",
          "content": "<p>but does it blend ? I have also a nn with .979x, it adds almost 0 to the ensemble. lgb ist at .981x. I feel that it must be tuned to equal levels like lgb, but this seems impossible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316953,
          "author_name": "hoogleraar",
          "author_url": "",
          "post_date": "04/20/2018 08:45:25",
          "content": "<p>you are right, it adds very tiny to the ensemble (~.0001) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316984,
          "author_name": "samehif",
          "author_url": "",
          "post_date": "04/20/2018 10:47:58",
          "content": "<p>Thanks Michael. is lbm same as Field wise Factorization Machine ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316986,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/20/2018 10:50:17",
          "content": "<p>I think it is lightgbm rather ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316988,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "04/20/2018 11:05:20",
          "content": "<p>Thanks @Michael for the info. If your NN is around 0.979, I won't even try NNs :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316993,
          "author_name": "mjahrer",
          "author_url": "",
          "post_date": "04/20/2018 11:38:04",
          "content": "<p>lbm  -&gt;  I mean lgb (light gbm) sorry!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318430,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "04/23/2018 20:26:17",
          "content": "<blockquote>\n  <p><strong>Michael Jahrer wrote</strong></p>\n  \n  <blockquote>\n    <p>but does it blend ? I have also a nn with .979x, it adds almost 0 to the ensemble. lgb ist at .981x. I feel that it must be tuned to equal levels like lgb, but this seems impossible.</p>\n  </blockquote>\n</blockquote>\n\n<p>Have you tried using <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53689\">Generative Adversarial Denoising Autoencoder</a>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318452,
          "author_name": "mjahrer",
          "author_url": "",
          "post_date": "04/23/2018 21:00:08",
          "content": "<p>haha, no, you ?</p>\n\n<p>but yea, I spend some days in re-writing DAE code to make this data transformation on-the-fly (otherwise mem would explode). very very slow here, 1 epoch through 242M samples with a 1100-2000-2000-1100 net is at 2700[s]. No way to effectively tune something. Also I feel this sparse data is not suitable for DAE. supervised nn afterwards scored also not good, doesnt help in blend. Probably settingd/params are in wrong range, I see no light here :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318454,
          "author_name": "rhgrossm",
          "author_url": "",
          "post_date": "04/23/2018 21:03:24",
          "content": "<p>I tried gans and autoencoders (separately :p) and they were between .9799 and .9802, not really worth the effort.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319416,
          "author_name": "its7171",
          "author_url": "",
          "post_date": "04/26/2018 02:19:56",
          "content": "<p>Blend of 3 model (1 LGBM(0.9815) model and 2 NN(0.9815) models) got LB: 09819</p>\n\n<p>It's just weighted average. It seems work without autoencoders, :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319480,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/26/2018 06:48:04",
          "content": "<p>@CuteChibiko, you convinced me to start on NN, so far working with a single lgb, I will soon reach a wall I think given how people ahead of me are stalled.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 317553,
      "author_name": "ianchute",
      "author_url": "",
      "post_date": "04/21/2018 20:50:52",
      "content": "<p>I was able to achieve 0.9110 using a simple Naive Bayes model. \n(Which is far from \"performing well\" though, but hey, it's not tree based!)\n<a href=\"https://www.kaggle.com/ianchute/naive-bayes-model-on-eight-basic-features/notebook\">https://www.kaggle.com/ianchute/naive-bayes-model-on-eight-basic-features/notebook</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 318396,
      "author_name": "sijunhe9248",
      "author_url": "",
      "post_date": "04/23/2018 18:42:33",
      "content": "<p>Seems like the only other model that comes close to LightGBM is FM_FTRL. From our experience, it also blends well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 318413,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/23/2018 19:22:29",
          "content": "<p>CuteChibiko mentions a .9814 neural network (of unknown architecture) <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\">in this thread</a>, so it would seem there's still hope for deep learning. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318429,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/23/2018 20:16:05",
          "content": "<p>Thank you <a href=\"/sijunhe\">@sijunhe</a>! May I ask what kind of parameters could be tuned for better performance in FM_FTRL?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319036,
          "author_name": "sijunhe9248",
          "author_url": "",
          "post_date": "04/25/2018 04:44:41",
          "content": "<p>Thanks for the heads up, Joe.</p>\n\n<p>@Snorlax, our FM_FTRL is about the same as the public kernel (~0.977). I didn't mean that FM_FTRL is as good as LGB at 0.981x., but it seemed like the model in public kernels with comparable performance. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 319204,
      "author_name": "marrvolo",
      "author_url": "",
      "post_date": "04/25/2018 13:52:10",
      "content": "<p>your single model lgb is at .981x, wow that's great, seems I should work harder :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "316759": "i wonder how far you can go with other model types besdie lgb/xgb. so far i have only tried lgb and it is performing really well and relatively fast for such a big data.",
    "316872": "NN can also perform well (.977-.978)",
    "316936": "but does it blend ? I have also a nn with .979x, it adds almost 0 to the ensemble. lgb ist at .981x. I feel that it must be tuned to equal levels like lgb, but this seems impossible.",
    "316953": "you are right, it adds very tiny to the ensemble (~.0001)",
    "316984": "Thanks Michael. is lbm same as Field wise Factorization Machine ?",
    "316986": "I think it is lightgbm rather ;)",
    "316988": "Thanks @Michael for the info. If your NN is around 0.979, I won't even try NNs :)",
    "316993": "lbm  -&gt;  I mean lgb (light gbm) sorry!",
    "317553": "I was able to achieve 0.9110 using a simple Naive Bayes model. \n(Which is far from \"performing well\" though, but hey, it's not tree based!)\nhttps://www.kaggle.com/ianchute/naive-bayes-model-on-eight-basic-features/notebook",
    "318396": "Seems like the only other model that comes close to LightGBM is FM_FTRL. From our experience, it also blends well.",
    "318413": "CuteChibiko mentions a .9814 neural network (of unknown architecture) [in this thread][1], so it would seem there's still hope for deep learning. \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634",
    "318429": "Thank you @sijunhe! May I ask what kind of parameters could be tuned for better performance in FM_FTRL?",
    "318430": "&gt; **Michael Jahrer wrote**\n&gt; \n&gt;&gt; but does it blend ? I have also a nn with .979x, it adds almost 0 to the ensemble. lgb ist at .981x. I feel that it must be tuned to equal levels like lgb, but this seems impossible.\n\nHave you tried using [Generative Adversarial Denoising Autoencoder][1]?\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53689",
    "318452": "haha, no, you ?\n\nbut yea, I spend some days in re-writing DAE code to make this data transformation on-the-fly (otherwise mem would explode). very very slow here, 1 epoch through 242M samples with a 1100-2000-2000-1100 net is at 2700[s]. No way to effectively tune something. Also I feel this sparse data is not suitable for DAE. supervised nn afterwards scored also not good, doesnt help in blend. Probably settingd/params are in wrong range, I see no light here :D",
    "318454": "I tried gans and autoencoders (separately :p) and they were between .9799 and .9802, not really worth the effort.",
    "319036": "Thanks for the heads up, Joe.\n\n@Snorlax, our FM_FTRL is about the same as the public kernel (~0.977). I didn't mean that FM_FTRL is as good as LGB at 0.981x., but it seemed like the model in public kernels with comparable performance.",
    "319204": "your single model lgb is at .981x, wow that's great, seems I should work harder :)",
    "319416": "Blend of 3 model (1 LGBM(0.9815) model and 2 NN(0.9815) models) got LB: 09819\n\nIt's just weighted average. It seems work without autoencoders, :)",
    "319480": "CuteChibiko, you convinced me to start on NN, so far working with a single lgb, I will soon reach a wall I think given how people ahead of me are stalled."
  },
  "source": "meta"
}