{
  "id": 71639,
  "title": "What are effective hyperparameters for RNN that are helpful?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71639",
  "author_name": "Neuron Engineer",
  "post_date": "2018-11-15T09:30:16.749000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I have been tried some hyperparameter search for RNN \n(I think everybody has done the same :D )</p>\n\n<p>Most of the attempts do not have any effects, only few seems help.</p>\n\n<p>Concretely, I've tried : (based on SRK RNN+Embedding codes)</p>\n\n<ul>\n<li>change between GRU &lt;--&gt; LSTM</li>\n<li>increase more layers</li>\n<li>increase/decrease latent dimension</li>\n<li>add more dropout</li>\n<li>add batch norm</li>\n<li>deeper dense</li>\n<li>add skip (residual) connection</li>\n<li>vary ensemble weights (weigths of each classifier)</li>\n<li>change max_len of sequence</li>\n<li>try oov_token</li>\n<li>post padding vs. pre padding</li>\n<li>reduce validation data number to increase more training data number</li>\n</ul>\n\n<p>None of the above looks helpful to me according to validation F1 (or maybe my implementation is not so good)</p>\n\n<p>Only few simple hyperparameters seem helpful:\n- number of epochs\n- batch size (a little help)</p>\n\n<p>Did you guys find the same situations? Or there are any other interesting hyperparameters to tune. Or there are more preprocessing technique?  I am currently try incoporating Dieter's great kernel of data cleaning, but still have to work more on it.</p>",
  "messages": [
    {
      "id": 421699,
      "postDate": "2018-11-15T09:30:16.750Z",
      "content": "<p>I have been tried some hyperparameter search for RNN \n(I think everybody has done the same :D )</p>\n\n<p>Most of the attempts do not have any effects, only few seems help.</p>\n\n<p>Concretely, I've tried : (based on SRK RNN+Embedding codes)</p>\n\n<ul>\n<li>change between GRU &lt;--&gt; LSTM</li>\n<li>increase more layers</li>\n<li>increase/decrease latent dimension</li>\n<li>add more dropout</li>\n<li>add batch norm</li>\n<li>deeper dense</li>\n<li>add skip (residual) connection</li>\n<li>vary ensemble weights (weigths of each classifier)</li>\n<li>change max_len of sequence</li>\n<li>try oov_token</li>\n<li>post padding vs. pre padding</li>\n<li>reduce validation data number to increase more training data number</li>\n</ul>\n\n<p>None of the above looks helpful to me according to validation F1 (or maybe my implementation is not so good)</p>\n\n<p>Only few simple hyperparameters seem helpful:\n- number of epochs\n- batch size (a little help)</p>\n\n<p>Did you guys find the same situations? Or there are any other interesting hyperparameters to tune. Or there are more preprocessing technique?  I am currently try incoporating Dieter's great kernel of data cleaning, but still have to work more on it.</p>",
      "rawMarkdown": "I have been tried some hyperparameter search for RNN \n(I think everybody has done the same :D )\n\nMost of the attempts do not have any effects, only few seems help.\n\nConcretely, I've tried : (based on SRK RNN+Embedding codes)\n\n- change between GRU &lt;--&gt; LSTM\n- increase more layers\n- increase/decrease latent dimension\n- add more dropout\n- add batch norm\n- deeper dense\n- add skip (residual) connection\n- vary ensemble weights (weigths of each classifier)\n- change max_len of sequence\n- try oov_token\n- post padding vs. pre padding\n- reduce validation data number to increase more training data number\n\nNone of the above looks helpful to me according to validation F1 (or maybe my implementation is not so good)\n\nOnly few simple hyperparameters seem helpful:\n- number of epochs\n- batch size (a little help)\n\nDid you guys find the same situations? Or there are any other interesting hyperparameters to tune. Or there are more preprocessing technique?  I am currently try incoporating Dieter's great kernel of data cleaning, but still have to work more on it.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "421699": "I have been tried some hyperparameter search for RNN \n(I think everybody has done the same :D )\n\nMost of the attempts do not have any effects, only few seems help.\n\nConcretely, I've tried : (based on SRK RNN+Embedding codes)\n\n- change between GRU &lt;--&gt; LSTM\n- increase more layers\n- increase/decrease latent dimension\n- add more dropout\n- add batch norm\n- deeper dense\n- add skip (residual) connection\n- vary ensemble weights (weigths of each classifier)\n- change max_len of sequence\n- try oov_token\n- post padding vs. pre padding\n- reduce validation data number to increase more training data number\n\nNone of the above looks helpful to me according to validation F1 (or maybe my implementation is not so good)\n\nOnly few simple hyperparameters seem helpful:\n- number of epochs\n- batch size (a little help)\n\nDid you guys find the same situations? Or there are any other interesting hyperparameters to tune. Or there are more preprocessing technique?  I am currently try incoporating Dieter's great kernel of data cleaning, but still have to work more on it."
  }
}