{
  "id": 63049,
  "title": "Inception V2 not learning. Ideas?",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/63049",
  "author_name": "",
  "post_date": "2018-08-11T01:42:21.804856Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm trying to train the Inception model from scratch as a learning exercise. I've tried running it for a week on 500,000 images in training set, and 50,000 validation set. Seems like a decent size. After a week, the training acuracy is ~90% and validation acuracy only gets to 70% before it starts to diverge. </p>\n\n<p>Does it look like there's anything off with my implementation (see <a href=\"https://github.com/formigone/tf-imagenet/blob/master/models/inception_v2_bn.py\">https://github.com/formigone/tf-imagenet/blob/master/models/inception_v2_bn.py</a>)? Any other advice? </p>\n\n<p>Thanks </p>",
  "messages": [
    {
      "id": "368837",
      "postDate": "08/11/2018 01:42:21",
      "content": "<p>I'm trying to train the Inception model from scratch as a learning exercise. I've tried running it for a week on 500,000 images in training set, and 50,000 validation set. Seems like a decent size. After a week, the training acuracy is ~90% and validation acuracy only gets to 70% before it starts to diverge. </p>\n\n<p>Does it look like there's anything off with my implementation (see <a href=\"https://github.com/formigone/tf-imagenet/blob/master/models/inception_v2_bn.py\">https://github.com/formigone/tf-imagenet/blob/master/models/inception_v2_bn.py</a>)? Any other advice? </p>\n\n<p>Thanks </p>",
      "rawMarkdown": "I'm trying to train the Inception model from scratch as a learning exercise. I've tried running it for a week on 500,000 images in training set, and 50,000 validation set. Seems like a decent size. After a week, the training acuracy is ~90% and validation acuracy only gets to 70% before it starts to diverge. \n\nDoes it look like there's anything off with my implementation (see https://github.com/formigone/tf-imagenet/blob/master/models/inception_v2_bn.py)? Any other advice? \n\nThanks",
      "votes": null
    },
    {
      "id": "368839",
      "postDate": "08/11/2018 02:19:24",
      "content": "<p>I haven't looked at the code, but this is a very difficult set. I have experienced similar results with other models. My guess is that the major problem is the number of classes. I am really a beginner in the field but i think that the optimisers find it hard to find the minimum of the cost function for so many classes. If i had more time on my hands i would probably try retraining with something that is more suitable for sparse dimensions. Again, i am no expert so this might be just nonsense rumblings. If anyone who knows more can comment it would definitely be helpful. </p>",
      "rawMarkdown": "I haven't looked at the code, but this is a very difficult set. I have experienced similar results with other models. My guess is that the major problem is the number of classes. I am really a beginner in the field but i think that the optimisers find it hard to find the minimum of the cost function for so many classes. If i had more time on my hands i would probably try retraining with something that is more suitable for sparse dimensions. Again, i am no expert so this might be just nonsense rumblings. If anyone who knows more can comment it would definitely be helpful.",
      "votes": null
    },
    {
      "id": "368845",
      "postDate": "08/11/2018 03:12:07",
      "content": "<p>I should have posted this with my original post, but this is what I've tried so far, all with more or less consistent results  (seemingly just overfitting the training set as described above):</p>\n\n<ul>\n<li>trained on a half dozen classes totalling ~6,000 training samples (all food classes)</li>\n<li>trained on 42 classes with a total of ~40,000 samples (classes include various reptiles and aquatic animals)</li>\n<li>added batch norm before every relu (which also means after every convolutional layer)</li>\n<li>dropout keep prob of 0.7 and 0.5</li>\n<li>input images are scaled/normalized with: <code>features / 255 - 0.5</code></li>\n<li>input images are randomly flipped horizontally or vertically,  with random hue adjustment (in a very small range). All these augmentations use native Tensorflow ops.</li>\n</ul>",
      "rawMarkdown": "I should have posted this with my original post, but this is what I've tried so far, all with more or less consistent results  (seemingly just overfitting the training set as described above):\n\n - trained on a half dozen classes totalling ~6,000 training samples (all food classes)\n - trained on 42 classes with a total of ~40,000 samples (classes include various reptiles and aquatic animals)\n - added batch norm before every relu (which also means after every convolutional layer)\n - dropout keep prob of 0.7 and 0.5\n - input images are scaled/normalized with: `features / 255 - 0.5`\n - input images are randomly flipped horizontally or vertically,  with random hue adjustment (in a very small range). All these augmentations use native Tensorflow ops.",
      "votes": null
    },
    {
      "id": "368883",
      "postDate": "08/11/2018 06:32:53",
      "content": "<p>Ok, you did much more than me...\nOne thing... Both training on food and on reptiles seems to train on a set with very little variance. I would probably choose 20 random classes and pick for each about 2000 samples, including another 2000 for none. Did you start with the zoo inception as base? </p>",
      "rawMarkdown": "Ok, you did much more than me...\nOne thing... Both training on food and on reptiles seems to train on a set with very little variance. I would probably choose 20 random classes and pick for each about 2000 samples, including another 2000 for none. Did you start with the zoo inception as base?",
      "votes": null
    },
    {
      "id": "377231",
      "postDate": "08/28/2018 19:19:59",
      "content": "<p>As an update: I was able to get my model to learn. Turns out that the problem was with my bad use Tensorflow 1.4 estimator API. I was running training an eval on the same process. So after so many steps of training, the evaluator would take over. After that, the training would resume, but the input pipeline would reset. Thus, I was only really training on the first few thousand samples over and over. That explains how the model was able to learn to a high level of accuracy so fast, and why it also quickly would stop generalizing across the validation set.</p>",
      "rawMarkdown": "As an update: I was able to get my model to learn. Turns out that the problem was with my bad use Tensorflow 1.4 estimator API. I was running training an eval on the same process. So after so many steps of training, the evaluator would take over. After that, the training would resume, but the input pipeline would reset. Thus, I was only really training on the first few thousand samples over and over. That explains how the model was able to learn to a high level of accuracy so fast, and why it also quickly would stop generalizing across the validation set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 368839,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "08/11/2018 02:19:24",
      "content": "<p>I haven't looked at the code, but this is a very difficult set. I have experienced similar results with other models. My guess is that the major problem is the number of classes. I am really a beginner in the field but i think that the optimisers find it hard to find the minimum of the cost function for so many classes. If i had more time on my hands i would probably try retraining with something that is more suitable for sparse dimensions. Again, i am no expert so this might be just nonsense rumblings. If anyone who knows more can comment it would definitely be helpful. </p>",
      "votes": null,
      "replies": [
        {
          "id": 368845,
          "author_name": "formigone",
          "author_url": "",
          "post_date": "08/11/2018 03:12:07",
          "content": "<p>I should have posted this with my original post, but this is what I've tried so far, all with more or less consistent results  (seemingly just overfitting the training set as described above):</p>\n\n<ul>\n<li>trained on a half dozen classes totalling ~6,000 training samples (all food classes)</li>\n<li>trained on 42 classes with a total of ~40,000 samples (classes include various reptiles and aquatic animals)</li>\n<li>added batch norm before every relu (which also means after every convolutional layer)</li>\n<li>dropout keep prob of 0.7 and 0.5</li>\n<li>input images are scaled/normalized with: <code>features / 255 - 0.5</code></li>\n<li>input images are randomly flipped horizontally or vertically,  with random hue adjustment (in a very small range). All these augmentations use native Tensorflow ops.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 368883,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "08/11/2018 06:32:53",
          "content": "<p>Ok, you did much more than me...\nOne thing... Both training on food and on reptiles seems to train on a set with very little variance. I would probably choose 20 random classes and pick for each about 2000 samples, including another 2000 for none. Did you start with the zoo inception as base? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 377231,
      "author_name": "formigone",
      "author_url": "",
      "post_date": "08/28/2018 19:19:59",
      "content": "<p>As an update: I was able to get my model to learn. Turns out that the problem was with my bad use Tensorflow 1.4 estimator API. I was running training an eval on the same process. So after so many steps of training, the evaluator would take over. After that, the training would resume, but the input pipeline would reset. Thus, I was only really training on the first few thousand samples over and over. That explains how the model was able to learn to a high level of accuracy so fast, and why it also quickly would stop generalizing across the validation set.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "368837": "I'm trying to train the Inception model from scratch as a learning exercise. I've tried running it for a week on 500,000 images in training set, and 50,000 validation set. Seems like a decent size. After a week, the training acuracy is ~90% and validation acuracy only gets to 70% before it starts to diverge. \n\nDoes it look like there's anything off with my implementation (see https://github.com/formigone/tf-imagenet/blob/master/models/inception_v2_bn.py)? Any other advice? \n\nThanks",
    "368839": "I haven't looked at the code, but this is a very difficult set. I have experienced similar results with other models. My guess is that the major problem is the number of classes. I am really a beginner in the field but i think that the optimisers find it hard to find the minimum of the cost function for so many classes. If i had more time on my hands i would probably try retraining with something that is more suitable for sparse dimensions. Again, i am no expert so this might be just nonsense rumblings. If anyone who knows more can comment it would definitely be helpful.",
    "368845": "I should have posted this with my original post, but this is what I've tried so far, all with more or less consistent results  (seemingly just overfitting the training set as described above):\n\n - trained on a half dozen classes totalling ~6,000 training samples (all food classes)\n - trained on 42 classes with a total of ~40,000 samples (classes include various reptiles and aquatic animals)\n - added batch norm before every relu (which also means after every convolutional layer)\n - dropout keep prob of 0.7 and 0.5\n - input images are scaled/normalized with: `features / 255 - 0.5`\n - input images are randomly flipped horizontally or vertically,  with random hue adjustment (in a very small range). All these augmentations use native Tensorflow ops.",
    "368883": "Ok, you did much more than me...\nOne thing... Both training on food and on reptiles seems to train on a set with very little variance. I would probably choose 20 random classes and pick for each about 2000 samples, including another 2000 for none. Did you start with the zoo inception as base?",
    "377231": "As an update: I was able to get my model to learn. Turns out that the problem was with my bad use Tensorflow 1.4 estimator API. I was running training an eval on the same process. So after so many steps of training, the evaluator would take over. After that, the training would resume, but the input pipeline would reset. Thus, I was only really training on the first few thousand samples over and over. That explains how the model was able to learn to a high level of accuracy so fast, and why it also quickly would stop generalizing across the validation set."
  },
  "source": "meta"
}