{
  "id": 90661,
  "title": "Memory Advantage of Using AdaDelta over RMSprop",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/90661",
  "author_name": "James Conley",
  "post_date": "2019-04-25T17:42:37.512000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I had been struggling using RMSprop because my CNN was only able to run with batches of around 6, and always crashed from running out of memory at a batch size of 12.  After switching to Adadelta I successfully ran with a batch size of 168, finding a sweet spot for training time to be around 129 (3 gpus so all batches are multiples of 3).  Over 10x improvement!</p>\n\n<p>I figured Adadelta would use less memory but I didn't expect such a difference!</p>\n\n<p>Has anyone else noticed this?</p>",
  "messages": [
    {
      "id": 523186,
      "postDate": "2019-04-25T17:42:37.513Z",
      "content": "<p>I had been struggling using RMSprop because my CNN was only able to run with batches of around 6, and always crashed from running out of memory at a batch size of 12.  After switching to Adadelta I successfully ran with a batch size of 168, finding a sweet spot for training time to be around 129 (3 gpus so all batches are multiples of 3).  Over 10x improvement!</p>\n\n<p>I figured Adadelta would use less memory but I didn't expect such a difference!</p>\n\n<p>Has anyone else noticed this?</p>",
      "rawMarkdown": "I had been struggling using RMSprop because my CNN was only able to run with batches of around 6, and always crashed from running out of memory at a batch size of 12.  After switching to Adadelta I successfully ran with a batch size of 168, finding a sweet spot for training time to be around 129 (3 gpus so all batches are multiples of 3).  Over 10x improvement!\n\nI figured Adadelta would use less memory but I didn't expect such a difference!\n\nHas anyone else noticed this?",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "523186": "I had been struggling using RMSprop because my CNN was only able to run with batches of around 6, and always crashed from running out of memory at a batch size of 12.  After switching to Adadelta I successfully ran with a batch size of 168, finding a sweet spot for training time to be around 129 (3 gpus so all batches are multiples of 3).  Over 10x improvement!\n\nI figured Adadelta would use less memory but I didn't expect such a difference!\n\nHas anyone else noticed this?"
  }
}