{
  "id": 39060,
  "title": "Running out of time?",
  "url": "/competitions/carvana-image-masking-challenge/discussion/39060",
  "author_name": "",
  "post_date": "2017-09-06T13:37:13.748351300Z",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>If people are interested in training models on large image inputs but are running out of time (or finding it difficult to get the networks to converge well) growing Unets works well - you can start by optimising your network on a smaller input size and then use the weights from this as a starting point for an increased input size. </p>\n\n<p>I worked from (512 x 512) -&gt; (1024 x 1024) -&gt; (1280 x 1280) -&gt; (1536 x 1536). It took about 50 epochs for the (512 x 512) to reach a minimum but from then on it only took 5 - 15 epochs for the others to converge. </p>\n\n<p>Working backwards through this process also looks like it has some merit (i.e. use the weights from the scaled up model as an alternative starting point for models with lower resolution inputs). So far this seems to add value on a CV set. Don't have a LB score for it yet, though.</p>",
  "messages": [
    {
      "id": "218953",
      "postDate": "09/06/2017 13:37:13",
      "content": "<p>If people are interested in training models on large image inputs but are running out of time (or finding it difficult to get the networks to converge well) growing Unets works well - you can start by optimising your network on a smaller input size and then use the weights from this as a starting point for an increased input size. </p>\n\n<p>I worked from (512 x 512) -&gt; (1024 x 1024) -&gt; (1280 x 1280) -&gt; (1536 x 1536). It took about 50 epochs for the (512 x 512) to reach a minimum but from then on it only took 5 - 15 epochs for the others to converge. </p>\n\n<p>Working backwards through this process also looks like it has some merit (i.e. use the weights from the scaled up model as an alternative starting point for models with lower resolution inputs). So far this seems to add value on a CV set. Don't have a LB score for it yet, though.</p>",
      "rawMarkdown": "If people are interested in training models on large image inputs but are running out of time (or finding it difficult to get the networks to converge well) growing Unets works well - you can start by optimising your network on a smaller input size and then use the weights from this as a starting point for an increased input size. \n\nI worked from (512 x 512) -&gt; (1024 x 1024) -&gt; (1280 x 1280) -&gt; (1536 x 1536). It took about 50 epochs for the (512 x 512) to reach a minimum but from then on it only took 5 - 15 epochs for the others to converge. \n\nWorking backwards through this process also looks like it has some merit (i.e. use the weights from the scaled up model as an alternative starting point for models with lower resolution inputs). So far this seems to add value on a CV set. Don't have a LB score for it yet, though.",
      "votes": null
    },
    {
      "id": "219122",
      "postDate": "09/07/2017 03:34:01",
      "content": "<p>Somewhat unrelated- is there actual benefit from having 1536x1536? Iirc the original dimensions were 1920x1280</p>",
      "rawMarkdown": "Somewhat unrelated- is there actual benefit from having 1536x1536? Iirc the original dimensions were 1920x1280",
      "votes": null
    },
    {
      "id": "219160",
      "postDate": "09/07/2017 07:07:05",
      "content": "<p>Not as much as the jump from 1024 -&gt; 1280, but I saw improvements from it, yes. It was a bit of an experiment to see what somewhat magnifying the original would do.</p>",
      "rawMarkdown": "Not as much as the jump from 1024 -&gt; 1280, but I saw improvements from it, yes. It was a bit of an experiment to see what somewhat magnifying the original would do.",
      "votes": null
    },
    {
      "id": "219171",
      "postDate": "09/07/2017 08:06:40",
      "content": "<p>Sounds interesting. In light of your topic title, my 1280^2 took almost a day to run, so the 2 days it would take to run the 1536^2 it is somewhat terrifying.</p>",
      "rawMarkdown": "Sounds interesting. In light of your topic title, my 1280^2 took almost a day to run, so the 2 days it would take to run the 1536^2 it is somewhat terrifying.",
      "votes": null
    },
    {
      "id": "219189",
      "postDate": "09/07/2017 08:49:52",
      "content": "<p>Yeah, that's what sent me off on the weights-sharing... It's difficult to test ideas when it takes such a colossal amount of time to run them, especially when you know a chunk of them aren't going to work anyway!</p>",
      "rawMarkdown": "Yeah, that's what sent me off on the weights-sharing... It's difficult to test ideas when it takes such a colossal amount of time to run them, especially when you know a chunk of them aren't going to work anyway!",
      "votes": null
    },
    {
      "id": "219208",
      "postDate": "09/07/2017 10:00:06",
      "content": "<p>fergusoci, do you use accumulate the gradients?\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125</a></p>",
      "rawMarkdown": "fergusoci, do you use accumulate the gradients?\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125",
      "votes": null
    },
    {
      "id": "219210",
      "postDate": "09/07/2017 10:08:36",
      "content": "<p>No, I haven't been. For the large image inputs I'm just using batch size = 1 but it's been ok without the gradient accumulating (possibly because the starting weights are sensible). I did experiment with accumulating gradients earlier on and it wasn't making a huge difference in my case.</p>",
      "rawMarkdown": "No, I haven't been. For the large image inputs I'm just using batch size = 1 but it's been ok without the gradient accumulating (possibly because the starting weights are sensible). I did experiment with accumulating gradients earlier on and it wasn't making a huge difference in my case.",
      "votes": null
    },
    {
      "id": "219296",
      "postDate": "09/07/2017 16:24:55",
      "content": "<p>Interestingly enough, the 1536^2 is relatively time efficient for the number of pixels used. It only takes about 30% more time per epoch, but has 44% more pixels than the 1280^2.</p>",
      "rawMarkdown": "Interestingly enough, the 1536^2 is relatively time efficient for the number of pixels used. It only takes about 30% more time per epoch, but has 44% more pixels than the 1280^2.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 219122,
      "author_name": "stevenknguyen",
      "author_url": "",
      "post_date": "09/07/2017 03:34:01",
      "content": "<p>Somewhat unrelated- is there actual benefit from having 1536x1536? Iirc the original dimensions were 1920x1280</p>",
      "votes": null,
      "replies": [
        {
          "id": 219160,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "09/07/2017 07:07:05",
          "content": "<p>Not as much as the jump from 1024 -&gt; 1280, but I saw improvements from it, yes. It was a bit of an experiment to see what somewhat magnifying the original would do.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 219171,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "09/07/2017 08:06:40",
          "content": "<p>Sounds interesting. In light of your topic title, my 1280^2 took almost a day to run, so the 2 days it would take to run the 1536^2 it is somewhat terrifying.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 219189,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "09/07/2017 08:49:52",
          "content": "<p>Yeah, that's what sent me off on the weights-sharing... It's difficult to test ideas when it takes such a colossal amount of time to run them, especially when you know a chunk of them aren't going to work anyway!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 219296,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "09/07/2017 16:24:55",
          "content": "<p>Interestingly enough, the 1536^2 is relatively time efficient for the number of pixels used. It only takes about 30% more time per epoch, but has 44% more pixels than the 1280^2.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 219208,
      "author_name": "markpopov",
      "author_url": "",
      "post_date": "09/07/2017 10:00:06",
      "content": "<p>fergusoci, do you use accumulate the gradients?\n<a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 219210,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "09/07/2017 10:08:36",
          "content": "<p>No, I haven't been. For the large image inputs I'm just using batch size = 1 but it's been ok without the gradient accumulating (possibly because the starting weights are sensible). I did experiment with accumulating gradients earlier on and it wasn't making a huge difference in my case.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "218953": "If people are interested in training models on large image inputs but are running out of time (or finding it difficult to get the networks to converge well) growing Unets works well - you can start by optimising your network on a smaller input size and then use the weights from this as a starting point for an increased input size. \n\nI worked from (512 x 512) -&gt; (1024 x 1024) -&gt; (1280 x 1280) -&gt; (1536 x 1536). It took about 50 epochs for the (512 x 512) to reach a minimum but from then on it only took 5 - 15 epochs for the others to converge. \n\nWorking backwards through this process also looks like it has some merit (i.e. use the weights from the scaled up model as an alternative starting point for models with lower resolution inputs). So far this seems to add value on a CV set. Don't have a LB score for it yet, though.",
    "219122": "Somewhat unrelated- is there actual benefit from having 1536x1536? Iirc the original dimensions were 1920x1280",
    "219160": "Not as much as the jump from 1024 -&gt; 1280, but I saw improvements from it, yes. It was a bit of an experiment to see what somewhat magnifying the original would do.",
    "219171": "Sounds interesting. In light of your topic title, my 1280^2 took almost a day to run, so the 2 days it would take to run the 1536^2 it is somewhat terrifying.",
    "219189": "Yeah, that's what sent me off on the weights-sharing... It's difficult to test ideas when it takes such a colossal amount of time to run them, especially when you know a chunk of them aren't going to work anyway!",
    "219208": "fergusoci, do you use accumulate the gradients?\nhttps://www.kaggle.com/c/carvana-image-masking-challenge/discussion/38125",
    "219210": "No, I haven't been. For the large image inputs I'm just using batch size = 1 but it's been ok without the gradient accumulating (possibly because the starting weights are sensible). I did experiment with accumulating gradients earlier on and it wasn't making a huge difference in my case.",
    "219296": "Interestingly enough, the 1536^2 is relatively time efficient for the number of pixels used. It only takes about 30% more time per epoch, but has 44% more pixels than the 1280^2."
  },
  "source": "meta"
}