{
  "id": 76906,
  "title": "Preferred optimizer?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76906",
  "author_name": "",
  "post_date": "2019-01-07T19:34:50.068119500Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Which optimization algorithm has given you the best results?\nI'm currently experimenting with SGD w/ nesterov momentum. </p>",
  "messages": [
    {
      "id": "451863",
      "postDate": "01/07/2019 19:34:50",
      "content": "<p>Which optimization algorithm has given you the best results?\nI'm currently experimenting with SGD w/ nesterov momentum. </p>",
      "rawMarkdown": "Which optimization algorithm has given you the best results?\nI'm currently experimenting with SGD w/ nesterov momentum.",
      "votes": null
    },
    {
      "id": "451995",
      "postDate": "01/08/2019 02:35:41",
      "content": "<p>Interesting academic paper that talks about using Adam to begin training and then switching to SGD w/ Nesterov later on.\n<a href=\"https://arxiv.org/abs/1712.07628\">https://arxiv.org/abs/1712.07628</a></p>",
      "rawMarkdown": "Interesting academic paper that talks about using Adam to begin training and then switching to SGD w/ Nesterov later on.\nhttps://arxiv.org/abs/1712.07628",
      "votes": null
    },
    {
      "id": "454186",
      "postDate": "01/11/2019 08:12:02",
      "content": "<p>Personally, I found the AdamW optimizer (which fixes Weight Decay Regularization) to be better in terms of convergence while also generalizing fairly well.</p>",
      "rawMarkdown": "Personally, I found the AdamW optimizer (which fixes Weight Decay Regularization) to be better in terms of convergence while also generalizing fairly well.",
      "votes": null
    },
    {
      "id": "454896",
      "postDate": "01/12/2019 12:58:21",
      "content": "<p>I tried AdamW but found its adding time to the processing. Do you not find it too slow and thus not a good choice to optimize the run speed here?</p>",
      "rawMarkdown": "I tried AdamW but found its adding time to the processing. Do you not find it too slow and thus not a good choice to optimize the run speed here?",
      "votes": null
    },
    {
      "id": "456456",
      "postDate": "01/15/2019 21:01:44",
      "content": "<p>May I know the implementation of adamW do you use?</p>",
      "rawMarkdown": "May I know the implementation of adamW do you use?",
      "votes": null
    },
    {
      "id": "475266",
      "postDate": "02/20/2019 14:41:32",
      "content": "<p>I managed to just squeeze in the optimizer by parallelizing the preprocessing tasks with dask. FYI, while I found that TensorFlow has an implementation in their contribution branch, it was GLambard’s implementation (<a href=\"https://github.com/GLambard/AdamW_Keras\">https://github.com/GLambard/AdamW_Keras</a>) of AdamW in Keras that led me to this score. So, props to him for sharing it.</p>\n\n<p>Even though, I could have used Keras’ TFOptimizer wrapper for TF Contib Optimizers, it just didn’t turn out to be compatible with my callbacks defined in Keras as they require the learning rate parameter to follow Keras’ naming conventions.</p>",
      "rawMarkdown": "I managed to just squeeze in the optimizer by parallelizing the preprocessing tasks with dask. FYI, while I found that TensorFlow has an implementation in their contribution branch, it was GLambard’s implementation (https://github.com/GLambard/AdamW_Keras) of AdamW in Keras that led me to this score. So, props to him for sharing it.\n\nEven though, I could have used Keras’ TFOptimizer wrapper for TF Contib Optimizers, it just didn’t turn out to be compatible with my callbacks defined in Keras as they require the learning rate parameter to follow Keras’ naming conventions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 451995,
      "author_name": "julius6",
      "author_url": "",
      "post_date": "01/08/2019 02:35:41",
      "content": "<p>Interesting academic paper that talks about using Adam to begin training and then switching to SGD w/ Nesterov later on.\n<a href=\"https://arxiv.org/abs/1712.07628\">https://arxiv.org/abs/1712.07628</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454186,
      "author_name": "ranik40",
      "author_url": "",
      "post_date": "01/11/2019 08:12:02",
      "content": "<p>Personally, I found the AdamW optimizer (which fixes Weight Decay Regularization) to be better in terms of convergence while also generalizing fairly well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454896,
          "author_name": "monduiz",
          "author_url": "",
          "post_date": "01/12/2019 12:58:21",
          "content": "<p>I tried AdamW but found its adding time to the processing. Do you not find it too slow and thus not a good choice to optimize the run speed here?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 456456,
          "author_name": "konohayui",
          "author_url": "",
          "post_date": "01/15/2019 21:01:44",
          "content": "<p>May I know the implementation of adamW do you use?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 475266,
          "author_name": "ranik40",
          "author_url": "",
          "post_date": "02/20/2019 14:41:32",
          "content": "<p>I managed to just squeeze in the optimizer by parallelizing the preprocessing tasks with dask. FYI, while I found that TensorFlow has an implementation in their contribution branch, it was GLambard’s implementation (<a href=\"https://github.com/GLambard/AdamW_Keras\">https://github.com/GLambard/AdamW_Keras</a>) of AdamW in Keras that led me to this score. So, props to him for sharing it.</p>\n\n<p>Even though, I could have used Keras’ TFOptimizer wrapper for TF Contib Optimizers, it just didn’t turn out to be compatible with my callbacks defined in Keras as they require the learning rate parameter to follow Keras’ naming conventions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "451863": "Which optimization algorithm has given you the best results?\nI'm currently experimenting with SGD w/ nesterov momentum.",
    "451995": "Interesting academic paper that talks about using Adam to begin training and then switching to SGD w/ Nesterov later on.\nhttps://arxiv.org/abs/1712.07628",
    "454186": "Personally, I found the AdamW optimizer (which fixes Weight Decay Regularization) to be better in terms of convergence while also generalizing fairly well.",
    "454896": "I tried AdamW but found its adding time to the processing. Do you not find it too slow and thus not a good choice to optimize the run speed here?",
    "456456": "May I know the implementation of adamW do you use?",
    "475266": "I managed to just squeeze in the optimizer by parallelizing the preprocessing tasks with dask. FYI, while I found that TensorFlow has an implementation in their contribution branch, it was GLambard’s implementation (https://github.com/GLambard/AdamW_Keras) of AdamW in Keras that led me to this score. So, props to him for sharing it.\n\nEven though, I could have used Keras’ TFOptimizer wrapper for TF Contib Optimizers, it just didn’t turn out to be compatible with my callbacks defined in Keras as they require the learning rate parameter to follow Keras’ naming conventions."
  },
  "source": "meta"
}