{
  "id": 51422,
  "title": "A few model suggestions",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51422",
  "author_name": "",
  "post_date": "2018-03-08T18:13:28.216533Z",
  "votes": 40,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In the past I have seen some algorithms perform particularly well for tasks like this:</p>\n\n<ol>\n<li><p>FFM (field aware factorization machines) - <a href=\"https://www.csie.ntu.edu.tw/~cjlin/libffm/\">https://www.csie.ntu.edu.tw/~cjlin/libffm/</a>. I feel like they have the potential to perform really well here. @rcarson also made a light-ffm - <a href=\"https://www.kaggle.com/c/outbrain-click-prediction/discussion/27892\">https://www.kaggle.com/c/outbrain-click-prediction/discussion/27892</a>.</p></li>\n<li><p>FTRL (follow the regularized leader) - also from outbrain I saw this model perform well - <a href=\"https://www.kaggle.com/sudalairajkumar/ftrl-starter-with-leakage-vars/comments\">https://www.kaggle.com/sudalairajkumar/ftrl-starter-with-leakage-vars/comments</a> (code from SRK - thanks! :))</p></li>\n<li><p>When I first looked at the data I thought - this might be a good data set to try out representation learning - which was a winning model from the Porto Seguro competition! @Michael Jahrer - provided tips and in excellent explanation for the model here - <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\">https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629</a>. After seeing him join this competition I feel like this further reaffirms my theory that I could do well here :) </p></li>\n</ol>",
  "messages": [
    {
      "id": "292831",
      "postDate": "03/08/2018 18:13:28",
      "content": "<p>In the past I have seen some algorithms perform particularly well for tasks like this:</p>\n\n<ol>\n<li><p>FFM (field aware factorization machines) - <a href=\"https://www.csie.ntu.edu.tw/~cjlin/libffm/\">https://www.csie.ntu.edu.tw/~cjlin/libffm/</a>. I feel like they have the potential to perform really well here. @rcarson also made a light-ffm - <a href=\"https://www.kaggle.com/c/outbrain-click-prediction/discussion/27892\">https://www.kaggle.com/c/outbrain-click-prediction/discussion/27892</a>.</p></li>\n<li><p>FTRL (follow the regularized leader) - also from outbrain I saw this model perform well - <a href=\"https://www.kaggle.com/sudalairajkumar/ftrl-starter-with-leakage-vars/comments\">https://www.kaggle.com/sudalairajkumar/ftrl-starter-with-leakage-vars/comments</a> (code from SRK - thanks! :))</p></li>\n<li><p>When I first looked at the data I thought - this might be a good data set to try out representation learning - which was a winning model from the Porto Seguro competition! @Michael Jahrer - provided tips and in excellent explanation for the model here - <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629\">https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629</a>. After seeing him join this competition I feel like this further reaffirms my theory that I could do well here :) </p></li>\n</ol>",
      "rawMarkdown": "In the past I have seen some algorithms perform particularly well for tasks like this:\n\n1. FFM (field aware factorization machines) - https://www.csie.ntu.edu.tw/~cjlin/libffm/. I feel like they have the potential to perform really well here. @rcarson also made a light-ffm - https://www.kaggle.com/c/outbrain-click-prediction/discussion/27892.\n\n2. FTRL (follow the regularized leader) - also from outbrain I saw this model perform well - https://www.kaggle.com/sudalairajkumar/ftrl-starter-with-leakage-vars/comments (code from SRK - thanks! :))\n\n3. When I first looked at the data I thought - this might be a good data set to try out representation learning - which was a winning model from the Porto Seguro competition! @Michael Jahrer - provided tips and in excellent explanation for the model here - https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629. After seeing him join this competition I feel like this further reaffirms my theory that I could do well here :)",
      "votes": null
    },
    {
      "id": "292891",
      "postDate": "03/08/2018 20:59:29",
      "content": "<p>The competition seems similar to Porto in that the target is binary and highly imbalanced. Also in that the sponsors are (or profess to be) more interested in an abstraction (safe driver, fraud) related to negative targets than in the target itself. As in Porto, it's not so much a matter of classifying cases according to likely target value as trying to ascertain something about the cases that will on average be indirectly reflected in the target value. It differs from Porto in that there are more cases and fewer raw features. But also the target is considerably more imbalanced, so we need the more cases.</p>",
      "rawMarkdown": "The competition seems similar to Porto in that the target is binary and highly imbalanced. Also in that the sponsors are (or profess to be) more interested in an abstraction (safe driver, fraud) related to negative targets than in the target itself. As in Porto, it's not so much a matter of classifying cases according to likely target value as trying to ascertain something about the cases that will on average be indirectly reflected in the target value. It differs from Porto in that there are more cases and fewer raw features. But also the target is considerably more imbalanced, so we need the more cases.",
      "votes": null
    },
    {
      "id": "292905",
      "postDate": "03/08/2018 21:19:51",
      "content": "<p>Another similarity I think is interesting is the ambiguity around treating features as numeric vs. categorical. We're told the features are categorical here, but from what I can tell there is a ton of signal in the ordering of the raw IPs (on the 100k subset, you can get a .7+ out of sample AUC with a logistic regression on raw IP). In seguro some of my models worked better when including both original numeric features and categorical encoded versions, I'd bet the same thing will be true here.</p>",
      "rawMarkdown": "Another similarity I think is interesting is the ambiguity around treating features as numeric vs. categorical. We're told the features are categorical here, but from what I can tell there is a ton of signal in the ordering of the raw IPs (on the 100k subset, you can get a .7+ out of sample AUC with a logistic regression on raw IP). In seguro some of my models worked better when including both original numeric features and categorical encoded versions, I'd bet the same thing will be true here.",
      "votes": null
    },
    {
      "id": "293701",
      "postDate": "03/10/2018 11:53:14",
      "content": "<p>Thank you for sharing this observation. One can learn so much from past competitions... </p>",
      "rawMarkdown": "Thank you for sharing this observation. One can learn so much from past competitions...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 292891,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/08/2018 20:59:29",
      "content": "<p>The competition seems similar to Porto in that the target is binary and highly imbalanced. Also in that the sponsors are (or profess to be) more interested in an abstraction (safe driver, fraud) related to negative targets than in the target itself. As in Porto, it's not so much a matter of classifying cases according to likely target value as trying to ascertain something about the cases that will on average be indirectly reflected in the target value. It differs from Porto in that there are more cases and fewer raw features. But also the target is considerably more imbalanced, so we need the more cases.</p>",
      "votes": null,
      "replies": [
        {
          "id": 292905,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "03/08/2018 21:19:51",
          "content": "<p>Another similarity I think is interesting is the ambiguity around treating features as numeric vs. categorical. We're told the features are categorical here, but from what I can tell there is a ton of signal in the ordering of the raw IPs (on the 100k subset, you can get a .7+ out of sample AUC with a logistic regression on raw IP). In seguro some of my models worked better when including both original numeric features and categorical encoded versions, I'd bet the same thing will be true here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293701,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/10/2018 11:53:14",
          "content": "<p>Thank you for sharing this observation. One can learn so much from past competitions... </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "292831": "In the past I have seen some algorithms perform particularly well for tasks like this:\n\n1. FFM (field aware factorization machines) - https://www.csie.ntu.edu.tw/~cjlin/libffm/. I feel like they have the potential to perform really well here. @rcarson also made a light-ffm - https://www.kaggle.com/c/outbrain-click-prediction/discussion/27892.\n\n2. FTRL (follow the regularized leader) - also from outbrain I saw this model perform well - https://www.kaggle.com/sudalairajkumar/ftrl-starter-with-leakage-vars/comments (code from SRK - thanks! :))\n\n3. When I first looked at the data I thought - this might be a good data set to try out representation learning - which was a winning model from the Porto Seguro competition! @Michael Jahrer - provided tips and in excellent explanation for the model here - https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44629. After seeing him join this competition I feel like this further reaffirms my theory that I could do well here :)",
    "292891": "The competition seems similar to Porto in that the target is binary and highly imbalanced. Also in that the sponsors are (or profess to be) more interested in an abstraction (safe driver, fraud) related to negative targets than in the target itself. As in Porto, it's not so much a matter of classifying cases according to likely target value as trying to ascertain something about the cases that will on average be indirectly reflected in the target value. It differs from Porto in that there are more cases and fewer raw features. But also the target is considerably more imbalanced, so we need the more cases.",
    "292905": "Another similarity I think is interesting is the ambiguity around treating features as numeric vs. categorical. We're told the features are categorical here, but from what I can tell there is a ton of signal in the ordering of the raw IPs (on the 100k subset, you can get a .7+ out of sample AUC with a logistic regression on raw IP). In seguro some of my models worked better when including both original numeric features and categorical encoded versions, I'd bet the same thing will be true here.",
    "293701": "Thank you for sharing this observation. One can learn so much from past competitions..."
  },
  "source": "meta"
}