{
  "id": 122044,
  "title": "3rd Place Solution",
  "url": "/competitions/vehicle/writeups/tauteam13-3rd-place-solution",
  "author_name": "",
  "post_date": "2019-12-18T13:01:03.097Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p><strong>INTRO</strong>\n100+ hours used for this project and having some fun. Basically, I did not know anything about machine learning and neural networks before September when the introductory course by Joni Kämäräinen started. But the topic is very interesting and that is why I spent quite some time to learn things. I am basically quite hardware oriented, so I like to implement or prototype something, not just look at what happens on the screen, but this was fun. Also, I have some kind of optimization gene (basic need to optimize everything), so this was very addictive to try to improve the accuracy gradually.</p>\n\n<p>I started the project and learning process by building my own CNN models. Testing how adding different layers affect on validation loss. Validation loss was my primary score during training. Testing some residuals etc. Own models were prone to overfitting, but batch normalization after each convolution helped. Also, I did not use yet image augmentation with my own models. Could not reach 80% accuracy and kept in mind that I can not create competitive CNN model without deep knowledge so I switched to pretrained models offered by Keras.</p>\n\n<p><strong>MODEL PARAMETERS</strong>\nFirst trained some small models and with smaller image sizes (64x64, 128x128), like Mobilenet(alpha=0.25) to achieve fast training to quickly test how model parameters like pooling affects. Testing also that how many Dense layers should be added on top of the model.</p>\n\n<p>Conclusion was:\n- pooling='avg'\n- Use dense layer size of the average pooling layer output, then add the final dense output layer\n- Actually, used in final trainings also one more dense between those previous, size of half of the first, not really confirmed is it any better, but got that feeling somehow</p>\n\n<p>Example: \nglobal_average_pooling2d_1   (None, 512) (last layer of the Keras model)\ndense_2 (Dense)              (None, 512) <br>\ndense_1 (Dense)              (None, 17)</p>\n\n<p>OR </p>\n\n<p>global_average_pooling2d_1   (None, 512) <br>\ndense_3 (Dense)              (None, 512) <br>\ndense_2 (Dense)              (None, 256) \ndense_1 (Dense)              (None, 17)</p>\n\n<p><strong>DATA AUGMENTATION AND TRAINING</strong>\nI started with Keras ImageDataGenerator. First flow from the memory, but it had some stablity issues. Then tried with flow_from_directory, it was very slow. But this is easy to use when you have limited memory in your computer. \nWell, because it was slow and I did not like the limited control of the image processing methods, I decided to use imgaug library. With imgaug you have nice control of the image augmentation. I think I used quite mild image processing, but many different processing sequentially. I mean I created many different augmented image sets and trained those all sequentially.\nFor example, I create first image set (from all train images) by adjusting brightness (Multiply(0.75, 1.25)) and cropping (CropAndPad(-0.1,0.1)), and train model with this only like 5 epochs. Then I create second image set like, flip horizontally some (Fliplr(0.5)) and add gaussian blur (GaussianBlur(sigma(0, 0.5))), and train again. Etc.</p>\n\n<p>If training accuracy tend to go higher than validation accuracy (or loss lower) I added more augmented image sets and decreased epochs for each set. In the end I had seven different image sets for training. One original and six augmented.\nNot really confirmed that is this better than basic augmentation by Keras, but it was my intuition and I made it this way.</p>\n\n<p>I tried also to use RandAugment, but it had some Tensorflow V1 functions in the code and did not work with Tensorflow V2, I did not have time to dig into that. It might be something to look at\n<a href=\"https://arxiv.org/abs/1909.13719\">https://arxiv.org/abs/1909.13719</a>\n<a href=\"https://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/autoaugment.py\">https://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/autoaugment.py</a></p>\n\n<p>I used 85% of images for training and 15% for validation. Stratified and different random_state for each model.</p>\n\n<p>After the val loss did not decrease anymore I finalized the model by training quickly with validation images. I left the original unprocessed images for validation and run the same sequential augmentation and training for the previous validation images. Only 2-3 epochs for each image set.</p>\n\n<p><strong>MODELS TRAINED</strong>\nSo, the target was to use several different CNNs and ensemble the probabilities for the final classification. Why use anything else than the best ones, so I used these accuracy tables to find the models for final training:\n<a href=\"https://github.com/keras-team/keras-applications\">https://github.com/keras-team/keras-applications</a>\n<a href=\"https://github.com/qubvel/classification_models\">https://github.com/qubvel/classification_models</a></p>\n\n<p>EfficientNet seemed to be the best, so that was the number one choice\n<a href=\"https://arxiv.org/abs/1905.11946\">https://arxiv.org/abs/1905.11946</a>\n<a href=\"https://github.com/qubvel/efficientnet\">https://github.com/qubvel/efficientnet</a></p>\n\n<p>I trained these for the final\nResNet152V2 (image size 224)\nInceptionV3 (224)\nInceptionResNetV2 (224)\nXception (224)\nNASNetMobile (224)\nEfficientNet-B3 (300)\nEfficientNet-B4 (380)</p>\n\n<p>Dropped out from the final\nMobileNetV2(alpha=1.4) (224)\nVGG19 (224)</p>\n\n<p>What I wanted to train also, but had some issues\nNASNetLarge (GPU memory insufficient)\nResNeXt101  (could not find library compatible with Keras API)\nEfficientNet-B5,B6,B7 (GPU memory insufficient)</p>\n\n<p>Scores:         val loss,   Kaggle Private (late submission tests)\nEfficientNetB3        0.2311       0.91213\nXception                  0.2478     0.90593\nInceptionResnetV2 0.2559    0.90643\nResNet152V2          0.2603   0.88950\nInceptionV3            0.2713     0.89151\nNASNet Mobile       0.2853    0.88749\nEfficientNetB4        NA          0.90895</p>\n\n<p>My computer was running on it's limits with EfficientNetB4 and my image augmentation, many freezes. So in the end I switched to Keras augmentation, I could not maybe use the full potential of B4. Missing also final training with validation data.</p>\n\n<p>Run out of time in the end so I was not able test many combinations, but these were the two best public before competition closed\nScore                           Private,    Public\nInceptionResnetV2 + Xception + EfficientNetB3 + EfficientNetB4  0.92588 0.92226\nInceptionResnetV2 + Xception + EfficientNetB3 + EfficientNetB4 + Inception 0.92437\n0.92226</p>\n\n<p>Some late submission tests,            Private,       Public\nInceptionResNetV2 + Xception + EfficientNetB3     0.92035    0.92026\nEfficientNetB3 + EfficientNetB4                                 0.91884     0.92176</p>\n\n<p><strong>OTHERS</strong>\n- Using only SGD optimizer from the beginning, can't remember why.\n- Using ModelCheckpoint and ReduceLROnPlateau callbacks.\n- Not using learning rate scheduler\n- Usually learning rate at the beginning 0.0005, dropping it manually if the start of the learning looks too rapid.\n- I wanted to train many of the models with 299x299 image size, but run out of time</p>\n\n<p>Tested also\n- class_weight in fit: not so good\n- quick test of final category weighting based on image count in each category: no improvement\n- Image zero padding before image resizing to keep the original aspect ratio: slight decrease in accuracy\n- Automated image selection for training, dropping out poor quality images. Using Laplacian() to detect image sharpness. Dropped out blurry and small images: not fully tested, but seems that no clear effect in accuracy. Maybe the blurry images even have some similar effect as dropouts inside CNN to create noise and thus prevent overfitting?\n- Pre-trained model layer freezing: not good</p>\n\n<p><strong>PROBLEMS</strong>\nMemory usage is high! This is solved of course with flow_from_directory, but it is slow, even with SSD.\nI have a Desktop with 32GB and 8GB GPU. 32GB was not enough for 224x224 training set, not even close! You need to set the Windows virtual memory very big. Memory usage was something like 70GB. For 300x300 I had to use for loops for casting to float32, otherwise memory run out after one or two cast. Python does not release memory correctly, in my opinion.\nFor 380x380 I split the training data half and used SSD to store the other half while training the other.</p>\n\n<p>Thanks! And sorry for my family as I was physically home, but not mentally present during this project. I will optimize this text later. Also could add link to github source.</p>",
  "messages": [
    {
      "id": "696979",
      "postDate": "12/17/2019 10:14:35",
      "content": "<p><strong>INTRO</strong>\n100+ hours used for this project and having some fun. Basically, I did not know anything about machine learning and neural networks before September when the introductory course by Joni Kämäräinen started. But the topic is very interesting and that is why I spent quite some time to learn things. I am basically quite hardware oriented, so I like to implement or prototype something, not just look at what happens on the screen, but this was fun. Also, I have some kind of optimization gene (basic need to optimize everything), so this was very addictive to try to improve the accuracy gradually.</p>\n\n<p>I started the project and learning process by building my own CNN models. Testing how adding different layers affect on validation loss. Validation loss was my primary score during training. Testing some residuals etc. Own models were prone to overfitting, but batch normalization after each convolution helped. Also, I did not use yet image augmentation with my own models. Could not reach 80% accuracy and kept in mind that I can not create competitive CNN model without deep knowledge so I switched to pretrained models offered by Keras.</p>\n\n<p><strong>MODEL PARAMETERS</strong>\nFirst trained some small models and with smaller image sizes (64x64, 128x128), like Mobilenet(alpha=0.25) to achieve fast training to quickly test how model parameters like pooling affects. Testing also that how many Dense layers should be added on top of the model.</p>\n\n<p>Conclusion was:\n- pooling='avg'\n- Use dense layer size of the average pooling layer output, then add the final dense output layer\n- Actually, used in final trainings also one more dense between those previous, size of half of the first, not really confirmed is it any better, but got that feeling somehow</p>\n\n<p>Example: \nglobal_average_pooling2d_1   (None, 512) (last layer of the Keras model)\ndense_2 (Dense)              (None, 512) <br>\ndense_1 (Dense)              (None, 17)</p>\n\n<p>OR </p>\n\n<p>global_average_pooling2d_1   (None, 512) <br>\ndense_3 (Dense)              (None, 512) <br>\ndense_2 (Dense)              (None, 256) \ndense_1 (Dense)              (None, 17)</p>\n\n<p><strong>DATA AUGMENTATION AND TRAINING</strong>\nI started with Keras ImageDataGenerator. First flow from the memory, but it had some stablity issues. Then tried with flow_from_directory, it was very slow. But this is easy to use when you have limited memory in your computer. \nWell, because it was slow and I did not like the limited control of the image processing methods, I decided to use imgaug library. With imgaug you have nice control of the image augmentation. I think I used quite mild image processing, but many different processing sequentially. I mean I created many different augmented image sets and trained those all sequentially.\nFor example, I create first image set (from all train images) by adjusting brightness (Multiply(0.75, 1.25)) and cropping (CropAndPad(-0.1,0.1)), and train model with this only like 5 epochs. Then I create second image set like, flip horizontally some (Fliplr(0.5)) and add gaussian blur (GaussianBlur(sigma(0, 0.5))), and train again. Etc.</p>\n\n<p>If training accuracy tend to go higher than validation accuracy (or loss lower) I added more augmented image sets and decreased epochs for each set. In the end I had seven different image sets for training. One original and six augmented.\nNot really confirmed that is this better than basic augmentation by Keras, but it was my intuition and I made it this way.</p>\n\n<p>I tried also to use RandAugment, but it had some Tensorflow V1 functions in the code and did not work with Tensorflow V2, I did not have time to dig into that. It might be something to look at\n<a href=\"https://arxiv.org/abs/1909.13719\">https://arxiv.org/abs/1909.13719</a>\n<a href=\"https://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/autoaugment.py\">https://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/autoaugment.py</a></p>\n\n<p>I used 85% of images for training and 15% for validation. Stratified and different random_state for each model.</p>\n\n<p>After the val loss did not decrease anymore I finalized the model by training quickly with validation images. I left the original unprocessed images for validation and run the same sequential augmentation and training for the previous validation images. Only 2-3 epochs for each image set.</p>\n\n<p><strong>MODELS TRAINED</strong>\nSo, the target was to use several different CNNs and ensemble the probabilities for the final classification. Why use anything else than the best ones, so I used these accuracy tables to find the models for final training:\n<a href=\"https://github.com/keras-team/keras-applications\">https://github.com/keras-team/keras-applications</a>\n<a href=\"https://github.com/qubvel/classification_models\">https://github.com/qubvel/classification_models</a></p>\n\n<p>EfficientNet seemed to be the best, so that was the number one choice\n<a href=\"https://arxiv.org/abs/1905.11946\">https://arxiv.org/abs/1905.11946</a>\n<a href=\"https://github.com/qubvel/efficientnet\">https://github.com/qubvel/efficientnet</a></p>\n\n<p>I trained these for the final\nResNet152V2 (image size 224)\nInceptionV3 (224)\nInceptionResNetV2 (224)\nXception (224)\nNASNetMobile (224)\nEfficientNet-B3 (300)\nEfficientNet-B4 (380)</p>\n\n<p>Dropped out from the final\nMobileNetV2(alpha=1.4) (224)\nVGG19 (224)</p>\n\n<p>What I wanted to train also, but had some issues\nNASNetLarge (GPU memory insufficient)\nResNeXt101  (could not find library compatible with Keras API)\nEfficientNet-B5,B6,B7 (GPU memory insufficient)</p>\n\n<p>Scores:         val loss,   Kaggle Private (late submission tests)\nEfficientNetB3        0.2311       0.91213\nXception                  0.2478     0.90593\nInceptionResnetV2 0.2559    0.90643\nResNet152V2          0.2603   0.88950\nInceptionV3            0.2713     0.89151\nNASNet Mobile       0.2853    0.88749\nEfficientNetB4        NA          0.90895</p>\n\n<p>My computer was running on it's limits with EfficientNetB4 and my image augmentation, many freezes. So in the end I switched to Keras augmentation, I could not maybe use the full potential of B4. Missing also final training with validation data.</p>\n\n<p>Run out of time in the end so I was not able test many combinations, but these were the two best public before competition closed\nScore                           Private,    Public\nInceptionResnetV2 + Xception + EfficientNetB3 + EfficientNetB4  0.92588 0.92226\nInceptionResnetV2 + Xception + EfficientNetB3 + EfficientNetB4 + Inception 0.92437\n0.92226</p>\n\n<p>Some late submission tests,            Private,       Public\nInceptionResNetV2 + Xception + EfficientNetB3     0.92035    0.92026\nEfficientNetB3 + EfficientNetB4                                 0.91884     0.92176</p>\n\n<p><strong>OTHERS</strong>\n- Using only SGD optimizer from the beginning, can't remember why.\n- Using ModelCheckpoint and ReduceLROnPlateau callbacks.\n- Not using learning rate scheduler\n- Usually learning rate at the beginning 0.0005, dropping it manually if the start of the learning looks too rapid.\n- I wanted to train many of the models with 299x299 image size, but run out of time</p>\n\n<p>Tested also\n- class_weight in fit: not so good\n- quick test of final category weighting based on image count in each category: no improvement\n- Image zero padding before image resizing to keep the original aspect ratio: slight decrease in accuracy\n- Automated image selection for training, dropping out poor quality images. Using Laplacian() to detect image sharpness. Dropped out blurry and small images: not fully tested, but seems that no clear effect in accuracy. Maybe the blurry images even have some similar effect as dropouts inside CNN to create noise and thus prevent overfitting?\n- Pre-trained model layer freezing: not good</p>\n\n<p><strong>PROBLEMS</strong>\nMemory usage is high! This is solved of course with flow_from_directory, but it is slow, even with SSD.\nI have a Desktop with 32GB and 8GB GPU. 32GB was not enough for 224x224 training set, not even close! You need to set the Windows virtual memory very big. Memory usage was something like 70GB. For 300x300 I had to use for loops for casting to float32, otherwise memory run out after one or two cast. Python does not release memory correctly, in my opinion.\nFor 380x380 I split the training data half and used SSD to store the other half while training the other.</p>\n\n<p>Thanks! And sorry for my family as I was physically home, but not mentally present during this project. I will optimize this text later. Also could add link to github source.</p>",
      "rawMarkdown": "**INTRO**\n100+ hours used for this project and having some fun. Basically, I did not know anything about machine learning and neural networks before September when the introductory course by Joni Kämäräinen started. But the topic is very interesting and that is why I spent quite some time to learn things. I am basically quite hardware oriented, so I like to implement or prototype something, not just look at what happens on the screen, but this was fun. Also, I have some kind of optimization gene (basic need to optimize everything), so this was very addictive to try to improve the accuracy gradually.\n\nI started the project and learning process by building my own CNN models. Testing how adding different layers affect on validation loss. Validation loss was my primary score during training. Testing some residuals etc. Own models were prone to overfitting, but batch normalization after each convolution helped. Also, I did not use yet image augmentation with my own models. Could not reach 80% accuracy and kept in mind that I can not create competitive CNN model without deep knowledge so I switched to pretrained models offered by Keras.\n\n**MODEL PARAMETERS**\nFirst trained some small models and with smaller image sizes (64x64, 128x128), like Mobilenet(alpha=0.25) to achieve fast training to quickly test how model parameters like pooling affects. Testing also that how many Dense layers should be added on top of the model.\n\nConclusion was:\n- pooling='avg'\n- Use dense layer size of the average pooling layer output, then add the final dense output layer\n- Actually, used in final trainings also one more dense between those previous, size of half of the first, not really confirmed is it any better, but got that feeling somehow\n\nExample: \nglobal_average_pooling2d_1   (None, 512) (last layer of the Keras model)\ndense_2 (Dense)              (None, 512)  \ndense_1 (Dense)              (None, 17)\n\nOR \n\nglobal_average_pooling2d_1   (None, 512)        \ndense_3 (Dense)              (None, 512)  \ndense_2 (Dense)              (None, 256) \ndense_1 (Dense)              (None, 17)\n\n\n**DATA AUGMENTATION AND TRAINING**\nI started with Keras ImageDataGenerator. First flow from the memory, but it had some stablity issues. Then tried with flow_from_directory, it was very slow. But this is easy to use when you have limited memory in your computer. \nWell, because it was slow and I did not like the limited control of the image processing methods, I decided to use imgaug library. With imgaug you have nice control of the image augmentation. I think I used quite mild image processing, but many different processing sequentially. I mean I created many different augmented image sets and trained those all sequentially.\nFor example, I create first image set (from all train images) by adjusting brightness (Multiply(0.75, 1.25)) and cropping (CropAndPad(-0.1,0.1)), and train model with this only like 5 epochs. Then I create second image set like, flip horizontally some (Fliplr(0.5)) and add gaussian blur (GaussianBlur(sigma(0, 0.5))), and train again. Etc.\n\nIf training accuracy tend to go higher than validation accuracy (or loss lower) I added more augmented image sets and decreased epochs for each set. In the end I had seven different image sets for training. One original and six augmented.\nNot really confirmed that is this better than basic augmentation by Keras, but it was my intuition and I made it this way.\n\nI tried also to use RandAugment, but it had some Tensorflow V1 functions in the code and did not work with Tensorflow V2, I did not have time to dig into that. It might be something to look at\nhttps://arxiv.org/abs/1909.13719\nhttps://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/autoaugment.py\n\nI used 85% of images for training and 15% for validation. Stratified and different random_state for each model.\n\nAfter the val loss did not decrease anymore I finalized the model by training quickly with validation images. I left the original unprocessed images for validation and run the same sequential augmentation and training for the previous validation images. Only 2-3 epochs for each image set.\n\n**MODELS TRAINED**\nSo, the target was to use several different CNNs and ensemble the probabilities for the final classification. Why use anything else than the best ones, so I used these accuracy tables to find the models for final training:\nhttps://github.com/keras-team/keras-applications\nhttps://github.com/qubvel/classification_models\n\nEfficientNet seemed to be the best, so that was the number one choice\nhttps://arxiv.org/abs/1905.11946\nhttps://github.com/qubvel/efficientnet\n\nI trained these for the final\nResNet152V2 (image size 224)\nInceptionV3 (224)\nInceptionResNetV2 (224)\nXception (224)\nNASNetMobile (224)\nEfficientNet-B3 (300)\nEfficientNet-B4 (380)\n\nDropped out from the final\nMobileNetV2(alpha=1.4) (224)\nVGG19 (224)\n\nWhat I wanted to train also, but had some issues\nNASNetLarge (GPU memory insufficient)\nResNeXt101  (could not find library compatible with Keras API)\nEfficientNet-B5,B6,B7 (GPU memory insufficient)\n\nScores:\t\t\tval loss,\tKaggle Private (late submission tests)\nEfficientNetB3        0.2311\t   0.91213\nXception                  0.2478     0.90593\nInceptionResnetV2 0.2559    0.90643\nResNet152V2          0.2603\t  0.88950\nInceptionV3            0.2713     0.89151\nNASNet Mobile       0.2853    0.88749\nEfficientNetB4        NA          0.90895\n\nMy computer was running on it's limits with EfficientNetB4 and my image augmentation, many freezes. So in the end I switched to Keras augmentation, I could not maybe use the full potential of B4. Missing also final training with validation data.\n\nRun out of time in the end so I was not able test many combinations, but these were the two best public before competition closed\nScore \t\t\t\t\t\t\tPrivate, \tPublic\nInceptionResnetV2 + Xception + EfficientNetB3 + EfficientNetB4 \t0.92588\t0.92226\nInceptionResnetV2 + Xception + EfficientNetB3 + EfficientNetB4 + Inception 0.92437\n0.92226\n\nSome late submission tests,            Private,       Public\nInceptionResNetV2 + Xception + EfficientNetB3     0.92035    0.92026\nEfficientNetB3 + EfficientNetB4                                 0.91884     0.92176\n\n**OTHERS**\n- Using only SGD optimizer from the beginning, can't remember why.\n- Using ModelCheckpoint and ReduceLROnPlateau callbacks.\n- Not using learning rate scheduler\n- Usually learning rate at the beginning 0.0005, dropping it manually if the start of the learning looks too rapid.\n- I wanted to train many of the models with 299x299 image size, but run out of time\n\nTested also\n- class_weight in fit: not so good\n- quick test of final category weighting based on image count in each category: no improvement\n- Image zero padding before image resizing to keep the original aspect ratio: slight decrease in accuracy\n- Automated image selection for training, dropping out poor quality images. Using Laplacian() to detect image sharpness. Dropped out blurry and small images: not fully tested, but seems that no clear effect in accuracy. Maybe the blurry images even have some similar effect as dropouts inside CNN to create noise and thus prevent overfitting?\n- Pre-trained model layer freezing: not good\n\n**PROBLEMS**\nMemory usage is high! This is solved of course with flow_from_directory, but it is slow, even with SSD.\nI have a Desktop with 32GB and 8GB GPU. 32GB was not enough for 224x224 training set, not even close! You need to set the Windows virtual memory very big. Memory usage was something like 70GB. For 300x300 I had to use for loops for casting to float32, otherwise memory run out after one or two cast. Python does not release memory correctly, in my opinion.\nFor 380x380 I split the training data half and used SSD to store the other half while training the other.\n\n\nThanks! And sorry for my family as I was physically home, but not mentally present during this project. I will optimize this text later. Also could add link to github source.",
      "votes": null
    },
    {
      "id": "697216",
      "postDate": "12/17/2019 16:17:34",
      "content": "<p>Congrats to the best TAU team, and thanks for the insightful description.</p>\n\n<p>Regarding the <code>flow_from_directory</code> function, it may help to use multithreading such that many threads feed data to the gpu. If SSD is the bottleneck, then resizing all images beforehand may speed up loading.</p>\n\n<p>Interesting to hear that so many were using efficientnet. I did not even know it until this course and it just popped as a side product while I shared some links.</p>\n\n<p>Congrats once more.</p>",
      "rawMarkdown": "Congrats to the best TAU team, and thanks for the insightful description.\n\nRegarding the `flow_from_directory` function, it may help to use multithreading such that many threads feed data to the gpu. If SSD is the bottleneck, then resizing all images beforehand may speed up loading.\n\nInteresting to hear that so many were using efficientnet. I did not even know it until this course and it just popped as a side product while I shared some links.\n\nCongrats once more.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 697216,
      "author_name": "mahehu",
      "author_url": "",
      "post_date": "12/17/2019 16:17:34",
      "content": "<p>Congrats to the best TAU team, and thanks for the insightful description.</p>\n\n<p>Regarding the <code>flow_from_directory</code> function, it may help to use multithreading such that many threads feed data to the gpu. If SSD is the bottleneck, then resizing all images beforehand may speed up loading.</p>\n\n<p>Interesting to hear that so many were using efficientnet. I did not even know it until this course and it just popped as a side product while I shared some links.</p>\n\n<p>Congrats once more.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "696979": "**INTRO**\n100+ hours used for this project and having some fun. Basically, I did not know anything about machine learning and neural networks before September when the introductory course by Joni Kämäräinen started. But the topic is very interesting and that is why I spent quite some time to learn things. I am basically quite hardware oriented, so I like to implement or prototype something, not just look at what happens on the screen, but this was fun. Also, I have some kind of optimization gene (basic need to optimize everything), so this was very addictive to try to improve the accuracy gradually.\n\nI started the project and learning process by building my own CNN models. Testing how adding different layers affect on validation loss. Validation loss was my primary score during training. Testing some residuals etc. Own models were prone to overfitting, but batch normalization after each convolution helped. Also, I did not use yet image augmentation with my own models. Could not reach 80% accuracy and kept in mind that I can not create competitive CNN model without deep knowledge so I switched to pretrained models offered by Keras.\n\n**MODEL PARAMETERS**\nFirst trained some small models and with smaller image sizes (64x64, 128x128), like Mobilenet(alpha=0.25) to achieve fast training to quickly test how model parameters like pooling affects. Testing also that how many Dense layers should be added on top of the model.\n\nConclusion was:\n- pooling='avg'\n- Use dense layer size of the average pooling layer output, then add the final dense output layer\n- Actually, used in final trainings also one more dense between those previous, size of half of the first, not really confirmed is it any better, but got that feeling somehow\n\nExample: \nglobal_average_pooling2d_1   (None, 512) (last layer of the Keras model)\ndense_2 (Dense)              (None, 512)  \ndense_1 (Dense)              (None, 17)\n\nOR \n\nglobal_average_pooling2d_1   (None, 512)        \ndense_3 (Dense)              (None, 512)  \ndense_2 (Dense)              (None, 256) \ndense_1 (Dense)              (None, 17)\n\n\n**DATA AUGMENTATION AND TRAINING**\nI started with Keras ImageDataGenerator. First flow from the memory, but it had some stablity issues. Then tried with flow_from_directory, it was very slow. But this is easy to use when you have limited memory in your computer. \nWell, because it was slow and I did not like the limited control of the image processing methods, I decided to use imgaug library. With imgaug you have nice control of the image augmentation. I think I used quite mild image processing, but many different processing sequentially. I mean I created many different augmented image sets and trained those all sequentially.\nFor example, I create first image set (from all train images) by adjusting brightness (Multiply(0.75, 1.25)) and cropping (CropAndPad(-0.1,0.1)), and train model with this only like 5 epochs. Then I create second image set like, flip horizontally some (Fliplr(0.5)) and add gaussian blur (GaussianBlur(sigma(0, 0.5))), and train again. Etc.\n\nIf training accuracy tend to go higher than validation accuracy (or loss lower) I added more augmented image sets and decreased epochs for each set. In the end I had seven different image sets for training. One original and six augmented.\nNot really confirmed that is this better than basic augmentation by Keras, but it was my intuition and I made it this way.\n\nI tried also to use RandAugment, but it had some Tensorflow V1 functions in the code and did not work with Tensorflow V2, I did not have time to dig into that. It might be something to look at\nhttps://arxiv.org/abs/1909.13719\nhttps://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/autoaugment.py\n\nI used 85% of images for training and 15% for validation. Stratified and different random_state for each model.\n\nAfter the val loss did not decrease anymore I finalized the model by training quickly with validation images. I left the original unprocessed images for validation and run the same sequential augmentation and training for the previous validation images. Only 2-3 epochs for each image set.\n\n**MODELS TRAINED**\nSo, the target was to use several different CNNs and ensemble the probabilities for the final classification. Why use anything else than the best ones, so I used these accuracy tables to find the models for final training:\nhttps://github.com/keras-team/keras-applications\nhttps://github.com/qubvel/classification_models\n\nEfficientNet seemed to be the best, so that was the number one choice\nhttps://arxiv.org/abs/1905.11946\nhttps://github.com/qubvel/efficientnet\n\nI trained these for the final\nResNet152V2 (image size 224)\nInceptionV3 (224)\nInceptionResNetV2 (224)\nXception (224)\nNASNetMobile (224)\nEfficientNet-B3 (300)\nEfficientNet-B4 (380)\n\nDropped out from the final\nMobileNetV2(alpha=1.4) (224)\nVGG19 (224)\n\nWhat I wanted to train also, but had some issues\nNASNetLarge (GPU memory insufficient)\nResNeXt101  (could not find library compatible with Keras API)\nEfficientNet-B5,B6,B7 (GPU memory insufficient)\n\nScores:\t\t\tval loss,\tKaggle Private (late submission tests)\nEfficientNetB3        0.2311\t   0.91213\nXception                  0.2478     0.90593\nInceptionResnetV2 0.2559    0.90643\nResNet152V2          0.2603\t  0.88950\nInceptionV3            0.2713     0.89151\nNASNet Mobile       0.2853    0.88749\nEfficientNetB4        NA          0.90895\n\nMy computer was running on it's limits with EfficientNetB4 and my image augmentation, many freezes. So in the end I switched to Keras augmentation, I could not maybe use the full potential of B4. Missing also final training with validation data.\n\nRun out of time in the end so I was not able test many combinations, but these were the two best public before competition closed\nScore \t\t\t\t\t\t\tPrivate, \tPublic\nInceptionResnetV2 + Xception + EfficientNetB3 + EfficientNetB4 \t0.92588\t0.92226\nInceptionResnetV2 + Xception + EfficientNetB3 + EfficientNetB4 + Inception 0.92437\n0.92226\n\nSome late submission tests,            Private,       Public\nInceptionResNetV2 + Xception + EfficientNetB3     0.92035    0.92026\nEfficientNetB3 + EfficientNetB4                                 0.91884     0.92176\n\n**OTHERS**\n- Using only SGD optimizer from the beginning, can't remember why.\n- Using ModelCheckpoint and ReduceLROnPlateau callbacks.\n- Not using learning rate scheduler\n- Usually learning rate at the beginning 0.0005, dropping it manually if the start of the learning looks too rapid.\n- I wanted to train many of the models with 299x299 image size, but run out of time\n\nTested also\n- class_weight in fit: not so good\n- quick test of final category weighting based on image count in each category: no improvement\n- Image zero padding before image resizing to keep the original aspect ratio: slight decrease in accuracy\n- Automated image selection for training, dropping out poor quality images. Using Laplacian() to detect image sharpness. Dropped out blurry and small images: not fully tested, but seems that no clear effect in accuracy. Maybe the blurry images even have some similar effect as dropouts inside CNN to create noise and thus prevent overfitting?\n- Pre-trained model layer freezing: not good\n\n**PROBLEMS**\nMemory usage is high! This is solved of course with flow_from_directory, but it is slow, even with SSD.\nI have a Desktop with 32GB and 8GB GPU. 32GB was not enough for 224x224 training set, not even close! You need to set the Windows virtual memory very big. Memory usage was something like 70GB. For 300x300 I had to use for loops for casting to float32, otherwise memory run out after one or two cast. Python does not release memory correctly, in my opinion.\nFor 380x380 I split the training data half and used SSD to store the other half while training the other.\n\n\nThanks! And sorry for my family as I was physically home, but not mentally present during this project. I will optimize this text later. Also could add link to github source.",
    "697216": "Congrats to the best TAU team, and thanks for the insightful description.\n\nRegarding the `flow_from_directory` function, it may help to use multithreading such that many threads feed data to the gpu. If SSD is the bottleneck, then resizing all images beforehand may speed up loading.\n\nInteresting to hear that so many were using efficientnet. I did not even know it until this course and it just popped as a side product while I shared some links.\n\nCongrats once more."
  },
  "source": "meta"
}