{
  "id": 53274,
  "title": "AutoEncoders, Anyone?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53274",
  "author_name": "",
  "post_date": "2018-03-28T21:25:26.680147700Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p><em>Warning: I am a newb.</em></p>\n\n<p>For the past couple of days I have been messing around with AutoEncoders for Anomaly Detection, since they can sometimes produce good results with imbalanced data. But I can't get past the AUC=0.7 mark on my validation sets.</p>\n\n<p>I use a simple architecture, with two Dense layers each for the Encoder/Decoder (pretty much <a href=\"https://medium.com/@curiousily/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd\">this</a>). I have used a couple of preprocessing schemes, like the one <a href=\"https://www.kaggle.com/joaopmpeinado/single-xgboost-lb-0-966/code\">here</a>.</p>\n\n<p>Has anyone gotten it to work? Any tips you can share?</p>",
  "messages": [
    {
      "id": "305433",
      "postDate": "03/28/2018 21:25:26",
      "content": "<p><em>Warning: I am a newb.</em></p>\n\n<p>For the past couple of days I have been messing around with AutoEncoders for Anomaly Detection, since they can sometimes produce good results with imbalanced data. But I can't get past the AUC=0.7 mark on my validation sets.</p>\n\n<p>I use a simple architecture, with two Dense layers each for the Encoder/Decoder (pretty much <a href=\"https://medium.com/@curiousily/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd\">this</a>). I have used a couple of preprocessing schemes, like the one <a href=\"https://www.kaggle.com/joaopmpeinado/single-xgboost-lb-0-966/code\">here</a>.</p>\n\n<p>Has anyone gotten it to work? Any tips you can share?</p>",
      "rawMarkdown": "*Warning: I am a newb.*\n\nFor the past couple of days I have been messing around with AutoEncoders for Anomaly Detection, since they can sometimes produce good results with imbalanced data. But I can't get past the AUC=0.7 mark on my validation sets.\n\nI use a simple architecture, with two Dense layers each for the Encoder/Decoder (pretty much [this](https://medium.com/@curiousily/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd)). I have used a couple of preprocessing schemes, like the one [here](https://www.kaggle.com/joaopmpeinado/single-xgboost-lb-0-966/code).\n\nHas anyone gotten it to work? Any tips you can share?",
      "votes": null
    },
    {
      "id": "305711",
      "postDate": "03/29/2018 09:55:31",
      "content": "<p>Given Michael Jahrer is active here yo may see a great use of autoencoders.  Check his winning solution in the Porto Seguro competition.</p>",
      "rawMarkdown": "Given Michael Jahrer is active here yo may see a great use of autoencoders.  Check his winning solution in the Porto Seguro competition.",
      "votes": null
    },
    {
      "id": "305728",
      "postDate": "03/29/2018 10:20:02",
      "content": "<p>I dont think AE will work here. They contribute best for dense data with moderate featCnt (10 .. 1000). Sparse data makes a problem because AE need to densify the data internally, coz it need to reconstruct all feats. And too many #rows leads to huge runtime, which makes tuning impractical. But lets see</p>",
      "rawMarkdown": "I dont think AE will work here. They contribute best for dense data with moderate featCnt (10 .. 1000). Sparse data makes a problem because AE need to densify the data internally, coz it need to reconstruct all feats. And too many #rows leads to huge runtime, which makes tuning impractical. But lets see",
      "votes": null
    },
    {
      "id": "305756",
      "postDate": "03/29/2018 10:55:18",
      "content": "<p>Thanks for the responses, much appreciated! I will let the AutoEncoder go for now, if someone can make it work, it's not me. :)</p>",
      "rawMarkdown": "Thanks for the responses, much appreciated! I will let the AutoEncoder go for now, if someone can make it work, it's not me. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 305711,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/29/2018 09:55:31",
      "content": "<p>Given Michael Jahrer is active here yo may see a great use of autoencoders.  Check his winning solution in the Porto Seguro competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 305728,
          "author_name": "mjahrer",
          "author_url": "",
          "post_date": "03/29/2018 10:20:02",
          "content": "<p>I dont think AE will work here. They contribute best for dense data with moderate featCnt (10 .. 1000). Sparse data makes a problem because AE need to densify the data internally, coz it need to reconstruct all feats. And too many #rows leads to huge runtime, which makes tuning impractical. But lets see</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 305756,
          "author_name": "antmarakis",
          "author_url": "",
          "post_date": "03/29/2018 10:55:18",
          "content": "<p>Thanks for the responses, much appreciated! I will let the AutoEncoder go for now, if someone can make it work, it's not me. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "305433": "*Warning: I am a newb.*\n\nFor the past couple of days I have been messing around with AutoEncoders for Anomaly Detection, since they can sometimes produce good results with imbalanced data. But I can't get past the AUC=0.7 mark on my validation sets.\n\nI use a simple architecture, with two Dense layers each for the Encoder/Decoder (pretty much [this](https://medium.com/@curiousily/credit-card-fraud-detection-using-autoencoders-in-keras-tensorflow-for-hackers-part-vii-20e0c85301bd)). I have used a couple of preprocessing schemes, like the one [here](https://www.kaggle.com/joaopmpeinado/single-xgboost-lb-0-966/code).\n\nHas anyone gotten it to work? Any tips you can share?",
    "305711": "Given Michael Jahrer is active here yo may see a great use of autoencoders.  Check his winning solution in the Porto Seguro competition.",
    "305728": "I dont think AE will work here. They contribute best for dense data with moderate featCnt (10 .. 1000). Sparse data makes a problem because AE need to densify the data internally, coz it need to reconstruct all feats. And too many #rows leads to huge runtime, which makes tuning impractical. But lets see",
    "305756": "Thanks for the responses, much appreciated! I will let the AutoEncoder go for now, if someone can make it work, it's not me. :)"
  },
  "source": "meta"
}