{
  "id": 102245,
  "title": "My First Competition",
  "url": "/competitions/recursion-cellular-image-classification/discussion/102245",
  "author_name": "",
  "post_date": "2019-07-31T19:39:11.788161300Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hey All!</p>\n\n<p>This is my first Kaggle competition. I picked it, because I figured training an image classification model should be pretty easy with the plethora of framework support for this application.</p>\n\n<p>In my life I have only trained some \"pre-canned\" things, like mnist examples, CIFAR10 examples, and also YOLOv2 (He gives you perfect step by step instructions).</p>\n\n<p>For this competition, I've chosen to use the Keras framework, and create a Sequential() model. For starters I'm just trying out the <a href=\"https://pjreddie.com/darknet/imagenet/#reference\">Darknet Reference Model</a>. I'm training on a GPU system with three P100.</p>\n\n<p>I'm struggling to gain confidence in my setup, because training seems to be so slow. Below is my log of 8 hours of training. I guess it is working, a bit, because the loss is decreasing, and the accuracy has already improved 40x compared to randomly guessing .0401 vs 0.001</p>\n\n<p>What do you guys think, am I on the right track? How much more patience should I have? Any words of advice? How can I gain more confidence in my approach? I'm terrified that I'll waste days of GPU time, and find I have some minuscule bug. One thing I did to try and gain confidence, is I used my infrastructure to train an MNIST model.</p>\n\n<p>Epoch 1/100\n2019-07-31 00:51:40.990750: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcublas.so.10.0\n2019-07-31 00:51:42.275405: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcudnn.so.7\n1939/1939 [==============================] - 5323s 3s/step - loss: 7.5341 - acc: 7.4136e-04 - val_loss: 7.1313 - val_acc: 9.1374e-04</p>\n\n<p>Epoch 00001: val_acc improved from -inf to 0.00091, saving model to my_model-01-0.00.h5\nEpoch 2/100\n1939/1939 [==============================] - 7501s 4s/step - loss: 7.1386 - acc: 7.8971e-04 - val_loss: 6.9789 - val_acc: 0.0010</p>\n\n<p>Epoch 00002: val_acc improved from 0.00091 to 0.00101, saving model to my_model-02-0.00.h5\nEpoch 3/100\n1939/1939 [==============================] - 7375s 4s/step - loss: 6.7877 - acc: 0.0027 - val_loss: 6.7479 - val_acc: 0.0044</p>\n\n<p>Epoch 00003: val_acc improved from 0.00101 to 0.00439, saving model to my_model-03-0.00.h5\nEpoch 4/100\n1939/1939 [==============================] - 7333s 4s/step - loss: 6.5184 - acc: 0.0058 - val_loss: 6.4502 - val_acc: 0.0078</p>\n\n<p>Epoch 00004: val_acc improved from 0.00439 to 0.00777, saving model to my_model-04-0.01.h5\nEpoch 5/100\n1939/1939 [==============================] - 6283s 3s/step - loss: 6.2537 - acc: 0.0125 - val_loss: 7.0372 - val_acc: 0.0069</p>\n\n<p>Epoch 00005: val_acc did not improve from 0.00777\nEpoch 6/100\n1939/1939 [==============================] - 7599s 4s/step - loss: 5.9548 - acc: 0.0237 - val_loss: 6.2632 - val_acc: 0.0178</p>\n\n<p>Epoch 00006: val_acc improved from 0.00777 to 0.01782, saving model to my_model-06-0.02.h5\nEpoch 7/100\n 220/1939 [==&gt;...........................] - ETA: 1:36:28 - loss: 5.6268 - acc: 0.0401</p>",
  "messages": [
    {
      "id": "589352",
      "postDate": "07/31/2019 19:39:11",
      "content": "<p>Hey All!</p>\n\n<p>This is my first Kaggle competition. I picked it, because I figured training an image classification model should be pretty easy with the plethora of framework support for this application.</p>\n\n<p>In my life I have only trained some \"pre-canned\" things, like mnist examples, CIFAR10 examples, and also YOLOv2 (He gives you perfect step by step instructions).</p>\n\n<p>For this competition, I've chosen to use the Keras framework, and create a Sequential() model. For starters I'm just trying out the <a href=\"https://pjreddie.com/darknet/imagenet/#reference\">Darknet Reference Model</a>. I'm training on a GPU system with three P100.</p>\n\n<p>I'm struggling to gain confidence in my setup, because training seems to be so slow. Below is my log of 8 hours of training. I guess it is working, a bit, because the loss is decreasing, and the accuracy has already improved 40x compared to randomly guessing .0401 vs 0.001</p>\n\n<p>What do you guys think, am I on the right track? How much more patience should I have? Any words of advice? How can I gain more confidence in my approach? I'm terrified that I'll waste days of GPU time, and find I have some minuscule bug. One thing I did to try and gain confidence, is I used my infrastructure to train an MNIST model.</p>\n\n<p>Epoch 1/100\n2019-07-31 00:51:40.990750: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcublas.so.10.0\n2019-07-31 00:51:42.275405: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcudnn.so.7\n1939/1939 [==============================] - 5323s 3s/step - loss: 7.5341 - acc: 7.4136e-04 - val_loss: 7.1313 - val_acc: 9.1374e-04</p>\n\n<p>Epoch 00001: val_acc improved from -inf to 0.00091, saving model to my_model-01-0.00.h5\nEpoch 2/100\n1939/1939 [==============================] - 7501s 4s/step - loss: 7.1386 - acc: 7.8971e-04 - val_loss: 6.9789 - val_acc: 0.0010</p>\n\n<p>Epoch 00002: val_acc improved from 0.00091 to 0.00101, saving model to my_model-02-0.00.h5\nEpoch 3/100\n1939/1939 [==============================] - 7375s 4s/step - loss: 6.7877 - acc: 0.0027 - val_loss: 6.7479 - val_acc: 0.0044</p>\n\n<p>Epoch 00003: val_acc improved from 0.00101 to 0.00439, saving model to my_model-03-0.00.h5\nEpoch 4/100\n1939/1939 [==============================] - 7333s 4s/step - loss: 6.5184 - acc: 0.0058 - val_loss: 6.4502 - val_acc: 0.0078</p>\n\n<p>Epoch 00004: val_acc improved from 0.00439 to 0.00777, saving model to my_model-04-0.01.h5\nEpoch 5/100\n1939/1939 [==============================] - 6283s 3s/step - loss: 6.2537 - acc: 0.0125 - val_loss: 7.0372 - val_acc: 0.0069</p>\n\n<p>Epoch 00005: val_acc did not improve from 0.00777\nEpoch 6/100\n1939/1939 [==============================] - 7599s 4s/step - loss: 5.9548 - acc: 0.0237 - val_loss: 6.2632 - val_acc: 0.0178</p>\n\n<p>Epoch 00006: val_acc improved from 0.00777 to 0.01782, saving model to my_model-06-0.02.h5\nEpoch 7/100\n 220/1939 [==&gt;...........................] - ETA: 1:36:28 - loss: 5.6268 - acc: 0.0401</p>",
      "rawMarkdown": "Hey All!\n\nThis is my first Kaggle competition. I picked it, because I figured training an image classification model should be pretty easy with the plethora of framework support for this application.\n\nIn my life I have only trained some \"pre-canned\" things, like mnist examples, CIFAR10 examples, and also YOLOv2 (He gives you perfect step by step instructions).\n\nFor this competition, I've chosen to use the Keras framework, and create a Sequential() model. For starters I'm just trying out the [Darknet Reference Model](https://pjreddie.com/darknet/imagenet/#reference). I'm training on a GPU system with three P100.\n\nI'm struggling to gain confidence in my setup, because training seems to be so slow. Below is my log of 8 hours of training. I guess it is working, a bit, because the loss is decreasing, and the accuracy has already improved 40x compared to randomly guessing .0401 vs 0.001\n\nWhat do you guys think, am I on the right track? How much more patience should I have? Any words of advice? How can I gain more confidence in my approach? I'm terrified that I'll waste days of GPU time, and find I have some minuscule bug. One thing I did to try and gain confidence, is I used my infrastructure to train an MNIST model.\n\n\nEpoch 1/100\n2019-07-31 00:51:40.990750: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcublas.so.10.0\n2019-07-31 00:51:42.275405: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcudnn.so.7\n1939/1939 [==============================] - 5323s 3s/step - loss: 7.5341 - acc: 7.4136e-04 - val_loss: 7.1313 - val_acc: 9.1374e-04\n\nEpoch 00001: val_acc improved from -inf to 0.00091, saving model to my_model-01-0.00.h5\nEpoch 2/100\n1939/1939 [==============================] - 7501s 4s/step - loss: 7.1386 - acc: 7.8971e-04 - val_loss: 6.9789 - val_acc: 0.0010\n\nEpoch 00002: val_acc improved from 0.00091 to 0.00101, saving model to my_model-02-0.00.h5\nEpoch 3/100\n1939/1939 [==============================] - 7375s 4s/step - loss: 6.7877 - acc: 0.0027 - val_loss: 6.7479 - val_acc: 0.0044\n\nEpoch 00003: val_acc improved from 0.00101 to 0.00439, saving model to my_model-03-0.00.h5\nEpoch 4/100\n1939/1939 [==============================] - 7333s 4s/step - loss: 6.5184 - acc: 0.0058 - val_loss: 6.4502 - val_acc: 0.0078\n\nEpoch 00004: val_acc improved from 0.00439 to 0.00777, saving model to my_model-04-0.01.h5\nEpoch 5/100\n1939/1939 [==============================] - 6283s 3s/step - loss: 6.2537 - acc: 0.0125 - val_loss: 7.0372 - val_acc: 0.0069\n\nEpoch 00005: val_acc did not improve from 0.00777\nEpoch 6/100\n1939/1939 [==============================] - 7599s 4s/step - loss: 5.9548 - acc: 0.0237 - val_loss: 6.2632 - val_acc: 0.0178\n\nEpoch 00006: val_acc improved from 0.00777 to 0.01782, saving model to my_model-06-0.02.h5\nEpoch 7/100\n 220/1939 [==&gt;...........................] - ETA: 1:36:28 - loss: 5.6268 - acc: 0.0401",
      "votes": null
    },
    {
      "id": "589561",
      "postDate": "08/01/2019 05:16:36",
      "content": "<p>Hello, my suggestion is to check your model in kaggle kernels, usually they are enough for 10-20 epochs. Have a good day!</p>",
      "rawMarkdown": "Hello, my suggestion is to check your model in kaggle kernels, usually they are enough for 10-20 epochs. Have a good day!",
      "votes": null
    },
    {
      "id": "589640",
      "postDate": "08/01/2019 07:29:20",
      "content": "<blockquote>\n  <p>How can I gain more confidence in my approach? </p>\n</blockquote>\n\n<p>I think <a href=\"http://karpathy.github.io/2019/04/25/recipe/\">http://karpathy.github.io/2019/04/25/recipe/</a> has a good approach for gaining confidence in one's pipeline.</p>",
      "rawMarkdown": "&gt; How can I gain more confidence in my approach? \n\nI think http://karpathy.github.io/2019/04/25/recipe/ has a good approach for gaining confidence in one's pipeline.",
      "votes": null
    },
    {
      "id": "589800",
      "postDate": "08/01/2019 11:42:36",
      "content": "<p>Костя, \nthanks a lot for sharing the link! The recommendations there are priceless!</p>",
      "rawMarkdown": "Костя, \nthanks a lot for sharing the link! The recommendations there are priceless!",
      "votes": null
    },
    {
      "id": "589819",
      "postDate": "08/01/2019 12:17:44",
      "content": "<p>wow 8 hours is a long time - have you tried a few different learning rates?</p>",
      "rawMarkdown": "wow 8 hours is a long time - have you tried a few different learning rates?",
      "votes": null
    },
    {
      "id": "589820",
      "postDate": "08/01/2019 12:18:22",
      "content": "<p>+100 for karpathy! I love his blog posts</p>",
      "rawMarkdown": "100 for karpathy! I love his blog posts",
      "votes": null
    },
    {
      "id": "591465",
      "postDate": "08/03/2019 18:11:42",
      "content": "<p>This is an awesome article! Thank you!</p>",
      "rawMarkdown": "This is an awesome article! Thank you!",
      "votes": null
    },
    {
      "id": "591466",
      "postDate": "08/03/2019 18:12:33",
      "content": "<p>I've just used the default learning rate for the Adam optimizer thus far.</p>",
      "rawMarkdown": "I've just used the default learning rate for the Adam optimizer thus far.",
      "votes": null
    },
    {
      "id": "591467",
      "postDate": "08/03/2019 18:14:38",
      "content": "<p>So after a few days of training, I am getting 99% on the training set, and 3% on the validation set. Need to look into why, and perhaps add dropout somewhere, maybe remove the FC at the end. Will go through the article shared by <a href=\"/lopuhin\">@lopuhin</a> first. Thanks to all!</p>",
      "rawMarkdown": "So after a few days of training, I am getting 99% on the training set, and 3% on the validation set. Need to look into why, and perhaps add dropout somewhere, maybe remove the FC at the end. Will go through the article shared by @lopuhin first. Thanks to all!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 589561,
      "author_name": "joven1997",
      "author_url": "",
      "post_date": "08/01/2019 05:16:36",
      "content": "<p>Hello, my suggestion is to check your model in kaggle kernels, usually they are enough for 10-20 epochs. Have a good day!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 589640,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "08/01/2019 07:29:20",
      "content": "<blockquote>\n  <p>How can I gain more confidence in my approach? </p>\n</blockquote>\n\n<p>I think <a href=\"http://karpathy.github.io/2019/04/25/recipe/\">http://karpathy.github.io/2019/04/25/recipe/</a> has a good approach for gaining confidence in one's pipeline.</p>",
      "votes": null,
      "replies": [
        {
          "id": 589800,
          "author_name": "samusram",
          "author_url": "",
          "post_date": "08/01/2019 11:42:36",
          "content": "<p>Костя, \nthanks a lot for sharing the link! The recommendations there are priceless!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 589820,
          "author_name": "hamishdickson",
          "author_url": "",
          "post_date": "08/01/2019 12:18:22",
          "content": "<p>+100 for karpathy! I love his blog posts</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 591465,
          "author_name": "wilderfield",
          "author_url": "",
          "post_date": "08/03/2019 18:11:42",
          "content": "<p>This is an awesome article! Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 589819,
      "author_name": "hamishdickson",
      "author_url": "",
      "post_date": "08/01/2019 12:17:44",
      "content": "<p>wow 8 hours is a long time - have you tried a few different learning rates?</p>",
      "votes": null,
      "replies": [
        {
          "id": 591466,
          "author_name": "wilderfield",
          "author_url": "",
          "post_date": "08/03/2019 18:12:33",
          "content": "<p>I've just used the default learning rate for the Adam optimizer thus far.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 591467,
      "author_name": "wilderfield",
      "author_url": "",
      "post_date": "08/03/2019 18:14:38",
      "content": "<p>So after a few days of training, I am getting 99% on the training set, and 3% on the validation set. Need to look into why, and perhaps add dropout somewhere, maybe remove the FC at the end. Will go through the article shared by <a href=\"/lopuhin\">@lopuhin</a> first. Thanks to all!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "589352": "Hey All!\n\nThis is my first Kaggle competition. I picked it, because I figured training an image classification model should be pretty easy with the plethora of framework support for this application.\n\nIn my life I have only trained some \"pre-canned\" things, like mnist examples, CIFAR10 examples, and also YOLOv2 (He gives you perfect step by step instructions).\n\nFor this competition, I've chosen to use the Keras framework, and create a Sequential() model. For starters I'm just trying out the [Darknet Reference Model](https://pjreddie.com/darknet/imagenet/#reference). I'm training on a GPU system with three P100.\n\nI'm struggling to gain confidence in my setup, because training seems to be so slow. Below is my log of 8 hours of training. I guess it is working, a bit, because the loss is decreasing, and the accuracy has already improved 40x compared to randomly guessing .0401 vs 0.001\n\nWhat do you guys think, am I on the right track? How much more patience should I have? Any words of advice? How can I gain more confidence in my approach? I'm terrified that I'll waste days of GPU time, and find I have some minuscule bug. One thing I did to try and gain confidence, is I used my infrastructure to train an MNIST model.\n\n\nEpoch 1/100\n2019-07-31 00:51:40.990750: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcublas.so.10.0\n2019-07-31 00:51:42.275405: I tensorflow/stream_executor/platform/default/dso_loader.cc:42] Successfully opened dynamic library libcudnn.so.7\n1939/1939 [==============================] - 5323s 3s/step - loss: 7.5341 - acc: 7.4136e-04 - val_loss: 7.1313 - val_acc: 9.1374e-04\n\nEpoch 00001: val_acc improved from -inf to 0.00091, saving model to my_model-01-0.00.h5\nEpoch 2/100\n1939/1939 [==============================] - 7501s 4s/step - loss: 7.1386 - acc: 7.8971e-04 - val_loss: 6.9789 - val_acc: 0.0010\n\nEpoch 00002: val_acc improved from 0.00091 to 0.00101, saving model to my_model-02-0.00.h5\nEpoch 3/100\n1939/1939 [==============================] - 7375s 4s/step - loss: 6.7877 - acc: 0.0027 - val_loss: 6.7479 - val_acc: 0.0044\n\nEpoch 00003: val_acc improved from 0.00101 to 0.00439, saving model to my_model-03-0.00.h5\nEpoch 4/100\n1939/1939 [==============================] - 7333s 4s/step - loss: 6.5184 - acc: 0.0058 - val_loss: 6.4502 - val_acc: 0.0078\n\nEpoch 00004: val_acc improved from 0.00439 to 0.00777, saving model to my_model-04-0.01.h5\nEpoch 5/100\n1939/1939 [==============================] - 6283s 3s/step - loss: 6.2537 - acc: 0.0125 - val_loss: 7.0372 - val_acc: 0.0069\n\nEpoch 00005: val_acc did not improve from 0.00777\nEpoch 6/100\n1939/1939 [==============================] - 7599s 4s/step - loss: 5.9548 - acc: 0.0237 - val_loss: 6.2632 - val_acc: 0.0178\n\nEpoch 00006: val_acc improved from 0.00777 to 0.01782, saving model to my_model-06-0.02.h5\nEpoch 7/100\n 220/1939 [==&gt;...........................] - ETA: 1:36:28 - loss: 5.6268 - acc: 0.0401",
    "589561": "Hello, my suggestion is to check your model in kaggle kernels, usually they are enough for 10-20 epochs. Have a good day!",
    "589640": "&gt; How can I gain more confidence in my approach? \n\nI think http://karpathy.github.io/2019/04/25/recipe/ has a good approach for gaining confidence in one's pipeline.",
    "589800": "Костя, \nthanks a lot for sharing the link! The recommendations there are priceless!",
    "589819": "wow 8 hours is a long time - have you tried a few different learning rates?",
    "589820": "100 for karpathy! I love his blog posts",
    "591465": "This is an awesome article! Thank you!",
    "591466": "I've just used the default learning rate for the Adam optimizer thus far.",
    "591467": "So after a few days of training, I am getting 99% on the training set, and 3% on the validation set. Need to look into why, and perhaps add dropout somewhere, maybe remove the FC at the end. Will go through the article shared by @lopuhin first. Thanks to all!"
  },
  "source": "meta"
}