{
  "id": 74530,
  "title": "Do I use autoencoders in a wrong manner?",
  "url": "/competitions/PLAsTiCC-2018/discussion/74530",
  "author_name": "",
  "post_date": "2018-12-13T07:02:19.177345Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>I tried to use a deep autoencoder as a feature extraction method. However, it didn't help me get good embeddings. The way I followed is below:</p>\n\n<p><img src=\"https://i.stack.imgur.com/yaPus.png\"></p>\n\n<p>1) Creating 2 datafarames for train_metadata and test_metadata. These dataframes contain handcrafted features which aren't used in my LGB model(meaning not useful for LGB)</p>\n\n<p>2) Merging 2 dataframes by pd.concat([train_metadata ,test_metadata],axis=0)</p>\n\n<p>3) Filling nans and infs with mean.</p>\n\n<p>4) Splitting into train and CV datasets(80:20).</p>\n\n<p>5) Using <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.MinMaxScaler.html\">sklearn</a>, I fitted MinMaxScaler on train data and transformed train and CV datasets.</p>\n\n<p>5) Using a deep autoencoder as an architecture. Number of nodes in layers are the following:\n265 ---&gt; 128 ---&gt; 32 ---&gt; 128 ---&gt; 265 . I used TensorFlow for this <a href=\"https://github.com/aymericdamien/TensorFlow-Examples/blob/master/examples/3_NeuralNetworks/autoencoder.py\">implementation</a>.</p>\n\n<p>6) I trained it. After training, I got embedded vector and used the embedded vector as new features for my LGB model. But, it worsened my LGB model.</p>\n\n<p>Can anyone tell me what is wrong?</p>\n\n<p>Thanks for your great contributions in advance.</p>",
  "messages": [
    {
      "id": "438142",
      "postDate": "12/13/2018 07:02:19",
      "content": "<p>Hello,</p>\n\n<p>I tried to use a deep autoencoder as a feature extraction method. However, it didn't help me get good embeddings. The way I followed is below:</p>\n\n<p><img src=\"https://i.stack.imgur.com/yaPus.png\"></p>\n\n<p>1) Creating 2 datafarames for train_metadata and test_metadata. These dataframes contain handcrafted features which aren't used in my LGB model(meaning not useful for LGB)</p>\n\n<p>2) Merging 2 dataframes by pd.concat([train_metadata ,test_metadata],axis=0)</p>\n\n<p>3) Filling nans and infs with mean.</p>\n\n<p>4) Splitting into train and CV datasets(80:20).</p>\n\n<p>5) Using <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.MinMaxScaler.html\">sklearn</a>, I fitted MinMaxScaler on train data and transformed train and CV datasets.</p>\n\n<p>5) Using a deep autoencoder as an architecture. Number of nodes in layers are the following:\n265 ---&gt; 128 ---&gt; 32 ---&gt; 128 ---&gt; 265 . I used TensorFlow for this <a href=\"https://github.com/aymericdamien/TensorFlow-Examples/blob/master/examples/3_NeuralNetworks/autoencoder.py\">implementation</a>.</p>\n\n<p>6) I trained it. After training, I got embedded vector and used the embedded vector as new features for my LGB model. But, it worsened my LGB model.</p>\n\n<p>Can anyone tell me what is wrong?</p>\n\n<p>Thanks for your great contributions in advance.</p>",
      "rawMarkdown": "Hello,\n\nI tried to use a deep autoencoder as a feature extraction method. However, it didn't help me get good embeddings. The way I followed is below:\n\n<img src=\"https://i.stack.imgur.com/yaPus.png\">\n\n1) Creating 2 datafarames for train_metadata and test_metadata. These dataframes contain handcrafted features which aren't used in my LGB model(meaning not useful for LGB)\n\n2) Merging 2 dataframes by pd.concat([train_metadata ,test_metadata],axis=0)\n\n3) Filling nans and infs with mean.\n\n4) Splitting into train and CV datasets(80:20).\n\n5) Using [sklearn](https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.MinMaxScaler.html), I fitted MinMaxScaler on train data and transformed train and CV datasets.\n\n5) Using a deep autoencoder as an architecture. Number of nodes in layers are the following:\n265 ---&gt; 128 ---&gt; 32 ---&gt; 128 ---&gt; 265 . I used TensorFlow for this [implementation](https://github.com/aymericdamien/TensorFlow-Examples/blob/master/examples/3_NeuralNetworks/autoencoder.py).\n\n6) I trained it. After training, I got embedded vector and used the embedded vector as new features for my LGB model. But, it worsened my LGB model.\n\nCan anyone tell me what is wrong?\n\nThanks for your great contributions in advance.",
      "votes": null
    },
    {
      "id": "438168",
      "postDate": "12/13/2018 07:53:13",
      "content": "<p>Autoencoding is not guaranteed to work as a feature extractor, especially with axis-aligned learners like LGB or other tree models. One of the issues with autoencoder is that you learn a dense, entangled (see <a href=\"https://arxiv.org/abs/1709.05047\">https://arxiv.org/abs/1709.05047</a> or <a href=\"https://www.youtube.com/watch?v=9zKuYvjFFS8\">this video</a> for some explanations) representation of the input space. This makes each individual feature less useful without the others, which is bad for LGB because it relies on each individual feature being meaningful on their own. You can try disentanged VAE described in the paper (you can find an implementation on github), or better, first examine if the features you feed into the network should contain useful information for classification.</p>\n\n<p>One potential case where autoencoder (and other dimensionality reduction techniques) may be useful is to serve as a non-linear PCA for highly correlated features.</p>",
      "rawMarkdown": "Autoencoding is not guaranteed to work as a feature extractor, especially with axis-aligned learners like LGB or other tree models. One of the issues with autoencoder is that you learn a dense, entangled (see https://arxiv.org/abs/1709.05047 or [this video](https://www.youtube.com/watch?v=9zKuYvjFFS8) for some explanations) representation of the input space. This makes each individual feature less useful without the others, which is bad for LGB because it relies on each individual feature being meaningful on their own. You can try disentanged VAE described in the paper (you can find an implementation on github), or better, first examine if the features you feed into the network should contain useful information for classification.\n\nOne potential case where autoencoder (and other dimensionality reduction techniques) may be useful is to serve as a non-linear PCA for highly correlated features.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 438168,
      "author_name": "mithrillion",
      "author_url": "",
      "post_date": "12/13/2018 07:53:13",
      "content": "<p>Autoencoding is not guaranteed to work as a feature extractor, especially with axis-aligned learners like LGB or other tree models. One of the issues with autoencoder is that you learn a dense, entangled (see <a href=\"https://arxiv.org/abs/1709.05047\">https://arxiv.org/abs/1709.05047</a> or <a href=\"https://www.youtube.com/watch?v=9zKuYvjFFS8\">this video</a> for some explanations) representation of the input space. This makes each individual feature less useful without the others, which is bad for LGB because it relies on each individual feature being meaningful on their own. You can try disentanged VAE described in the paper (you can find an implementation on github), or better, first examine if the features you feed into the network should contain useful information for classification.</p>\n\n<p>One potential case where autoencoder (and other dimensionality reduction techniques) may be useful is to serve as a non-linear PCA for highly correlated features.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "438142": "Hello,\n\nI tried to use a deep autoencoder as a feature extraction method. However, it didn't help me get good embeddings. The way I followed is below:\n\n<img src=\"https://i.stack.imgur.com/yaPus.png\">\n\n1) Creating 2 datafarames for train_metadata and test_metadata. These dataframes contain handcrafted features which aren't used in my LGB model(meaning not useful for LGB)\n\n2) Merging 2 dataframes by pd.concat([train_metadata ,test_metadata],axis=0)\n\n3) Filling nans and infs with mean.\n\n4) Splitting into train and CV datasets(80:20).\n\n5) Using [sklearn](https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.MinMaxScaler.html), I fitted MinMaxScaler on train data and transformed train and CV datasets.\n\n5) Using a deep autoencoder as an architecture. Number of nodes in layers are the following:\n265 ---&gt; 128 ---&gt; 32 ---&gt; 128 ---&gt; 265 . I used TensorFlow for this [implementation](https://github.com/aymericdamien/TensorFlow-Examples/blob/master/examples/3_NeuralNetworks/autoencoder.py).\n\n6) I trained it. After training, I got embedded vector and used the embedded vector as new features for my LGB model. But, it worsened my LGB model.\n\nCan anyone tell me what is wrong?\n\nThanks for your great contributions in advance.",
    "438168": "Autoencoding is not guaranteed to work as a feature extractor, especially with axis-aligned learners like LGB or other tree models. One of the issues with autoencoder is that you learn a dense, entangled (see https://arxiv.org/abs/1709.05047 or [this video](https://www.youtube.com/watch?v=9zKuYvjFFS8) for some explanations) representation of the input space. This makes each individual feature less useful without the others, which is bad for LGB because it relies on each individual feature being meaningful on their own. You can try disentanged VAE described in the paper (you can find an implementation on github), or better, first examine if the features you feed into the network should contain useful information for classification.\n\nOne potential case where autoencoder (and other dimensionality reduction techniques) may be useful is to serve as a non-linear PCA for highly correlated features."
  },
  "source": "meta"
}