{
  "id": 319061,
  "title": "A CNN is (not) All you Need [10th place]",
  "url": "/competitions/ultra-mnist/writeups/massimiliano-viola-a-cnn-is-not-all-you-need-10th-",
  "author_name": "",
  "post_date": "2022-04-15T08:14:20.741033800Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to the organizers for creating a really interesting challenge inspired by a class of real-world problems and with great prizes. If I am being completely honest, I joined the competition with the old data just to have fun with OpenCV, but as it immediately became clear that it just solved the task and the dataset changed, I stuck around nonetheless as this was a nice opportunity to learn a bit of PyTorch as a loyal TensorFlow user. </p>\n<h1>Validation scheme</h1>\n<p>Many reported that the test accuracy was lower than the validation one in the forum at the beginning, but after making sure train and test data came from the same distribution and submitting myself, I did not notice any issue. Since performing regular cross-validation was virtually impossible on large image sizes within the weekly GPU quota, I used a 10% hold-out set and it worked just fine.</p>\n<h1>Training methodology</h1>\n<p>After thinking for some days about how to approach the innovation track, I did not come up with a good idea so I just trained a bunch of CNNs in PyTorch with Kaggle and Colab GPUs. To speed up training, I tried the progressive resizing as described in the <a href=\"https://arxiv.org/pdf/2104.00298.pdf\" target=\"_blank\">EfficientNetV2 paper</a> and it worked just wonderfully. The dataset was noisy enough to prevent overfitting for a while, but in the end, I had to introduce some simple augmentation as ShiftScaleRotate, which was also increased during training alongside the image size.</p>\n<h1>Models</h1>\n<p>My solution is a pretty simple blend of my two top-scoring models, predicting the original dataset and the <a href=\"https://www.kaggle.com/competitions/ultra-mnist/discussion/313879\" target=\"_blank\">inverted colour one</a>. They are respectively an EfficientNetB3 and an EfficientNetV2B3 with image sizes 1024 for a total of four sets of predictions. The models all have CV/LB in the 0.8 range but are quite different. Using the hold-out dataset which contained the same image ids for all models, I found the best weights to combine these four predictions with Optuna and this gave me a decent boost, with my final submission scoring in the 0.85 range.</p>\n<h1>Key takeaways</h1>\n<ul>\n<li>CNNs are extremely powerful. Really, I would not have bet on such a high accuracy with such a noisy background and various digit sizes inside it, but I was totally wrong.</li>\n<li>Pytorch is great and extremely beginner-friendly. I will continue to learn it for sure because its flexibility is just priceless.</li>\n<li>Processing large image sizes is really time and memory consuming. This competition has brought to the fore situations where resizing is not an option if you want to achieve very high accuracy.</li>\n</ul>\n<p>Thanks for taking the time to read this. I really hope to see an innovation track solution as this is an important problem, but I also don't mind seeing how the winners broke the competition again. Congratulations to all participants, let's keep learning!</p>",
  "messages": [
    {
      "id": "1756072",
      "postDate": "04/15/2022 08:14:20",
      "content": "<p>Thanks to the organizers for creating a really interesting challenge inspired by a class of real-world problems and with great prizes. If I am being completely honest, I joined the competition with the old data just to have fun with OpenCV, but as it immediately became clear that it just solved the task and the dataset changed, I stuck around nonetheless as this was a nice opportunity to learn a bit of PyTorch as a loyal TensorFlow user. </p>\n<h1>Validation scheme</h1>\n<p>Many reported that the test accuracy was lower than the validation one in the forum at the beginning, but after making sure train and test data came from the same distribution and submitting myself, I did not notice any issue. Since performing regular cross-validation was virtually impossible on large image sizes within the weekly GPU quota, I used a 10% hold-out set and it worked just fine.</p>\n<h1>Training methodology</h1>\n<p>After thinking for some days about how to approach the innovation track, I did not come up with a good idea so I just trained a bunch of CNNs in PyTorch with Kaggle and Colab GPUs. To speed up training, I tried the progressive resizing as described in the <a href=\"https://arxiv.org/pdf/2104.00298.pdf\" target=\"_blank\">EfficientNetV2 paper</a> and it worked just wonderfully. The dataset was noisy enough to prevent overfitting for a while, but in the end, I had to introduce some simple augmentation as ShiftScaleRotate, which was also increased during training alongside the image size.</p>\n<h1>Models</h1>\n<p>My solution is a pretty simple blend of my two top-scoring models, predicting the original dataset and the <a href=\"https://www.kaggle.com/competitions/ultra-mnist/discussion/313879\" target=\"_blank\">inverted colour one</a>. They are respectively an EfficientNetB3 and an EfficientNetV2B3 with image sizes 1024 for a total of four sets of predictions. The models all have CV/LB in the 0.8 range but are quite different. Using the hold-out dataset which contained the same image ids for all models, I found the best weights to combine these four predictions with Optuna and this gave me a decent boost, with my final submission scoring in the 0.85 range.</p>\n<h1>Key takeaways</h1>\n<ul>\n<li>CNNs are extremely powerful. Really, I would not have bet on such a high accuracy with such a noisy background and various digit sizes inside it, but I was totally wrong.</li>\n<li>Pytorch is great and extremely beginner-friendly. I will continue to learn it for sure because its flexibility is just priceless.</li>\n<li>Processing large image sizes is really time and memory consuming. This competition has brought to the fore situations where resizing is not an option if you want to achieve very high accuracy.</li>\n</ul>\n<p>Thanks for taking the time to read this. I really hope to see an innovation track solution as this is an important problem, but I also don't mind seeing how the winners broke the competition again. Congratulations to all participants, let's keep learning!</p>",
      "rawMarkdown": "Thanks to the organizers for creating a really interesting challenge inspired by a class of real-world problems and with great prizes. If I am being completely honest, I joined the competition with the old data just to have fun with OpenCV, but as it immediately became clear that it just solved the task and the dataset changed, I stuck around nonetheless as this was a nice opportunity to learn a bit of PyTorch as a loyal TensorFlow user. \n\n# Validation scheme\nMany reported that the test accuracy was lower than the validation one in the forum at the beginning, but after making sure train and test data came from the same distribution and submitting myself, I did not notice any issue. Since performing regular cross-validation was virtually impossible on large image sizes within the weekly GPU quota, I used a 10% hold-out set and it worked just fine.\n\n# Training methodology\nAfter thinking for some days about how to approach the innovation track, I did not come up with a good idea so I just trained a bunch of CNNs in PyTorch with Kaggle and Colab GPUs. To speed up training, I tried the progressive resizing as described in the [EfficientNetV2 paper](https://arxiv.org/pdf/2104.00298.pdf) and it worked just wonderfully. The dataset was noisy enough to prevent overfitting for a while, but in the end, I had to introduce some simple augmentation as ShiftScaleRotate, which was also increased during training alongside the image size.\n\n# Models\nMy solution is a pretty simple blend of my two top-scoring models, predicting the original dataset and the [inverted colour one](https://www.kaggle.com/competitions/ultra-mnist/discussion/313879). They are respectively an EfficientNetB3 and an EfficientNetV2B3 with image sizes 1024 for a total of four sets of predictions. The models all have CV/LB in the 0.8 range but are quite different. Using the hold-out dataset which contained the same image ids for all models, I found the best weights to combine these four predictions with Optuna and this gave me a decent boost, with my final submission scoring in the 0.85 range.\n\n# Key takeaways\n- CNNs are extremely powerful. Really, I would not have bet on such a high accuracy with such a noisy background and various digit sizes inside it, but I was totally wrong.\n- Pytorch is great and extremely beginner-friendly. I will continue to learn it for sure because its flexibility is just priceless.\n- Processing large image sizes is really time and memory consuming. This competition has brought to the fore situations where resizing is not an option if you want to achieve very high accuracy.\n\nThanks for taking the time to read this. I really hope to see an innovation track solution as this is an important problem, but I also don't mind seeing how the winners broke the competition again. Congratulations to all participants, let's keep learning!",
      "votes": null
    },
    {
      "id": "1756448",
      "postDate": "04/15/2022 14:25:28",
      "content": "<p>Could you share the code?</p>",
      "rawMarkdown": "Could you share the code?",
      "votes": null
    },
    {
      "id": "1756523",
      "postDate": "04/15/2022 15:29:23",
      "content": "<p>Sorry, I don't feel very confident sharing the code right now as it is quite messy (I am learning PyTorch in these weeks) and also really similar to the <a href=\"https://www.kaggle.com/code/surajsharan/ultramnist-starter-notebook\" target=\"_blank\">baseline</a>. These are the main changes I've made to it:</p>\n<ul>\n<li>Used <a href=\"https://www.kaggle.com/datasets/kozodoi/timm-pytorch-image-models\" target=\"_blank\">timm</a> so I could have all models at hand</li>\n<li>Removed resizing and added ShiftScaleRotate from Albumentations to the train transformations</li>\n<li>Modified the training loop to keep track of train and validation metrics (accuracy and loss) so that I could plot their evolution in time later </li>\n<li>Put the model in eval mode when predicting valid data</li>\n<li>Added the possibility to save the class probabilities for the hold-out set and the test data after loading the best checkpoint</li>\n</ul>\n<p>I then resized and saved images in 384, 512, 768, and 1024 beforehand to speed things up. At each run of the notebook, I would load images of a certain size, change the intensity of augmentation and learning rate accordingly, then load the previous training checkpoint and resume from there. For example, first I trained on image size 384 for 15 epochs, then on image size 512 for 10 epochs, next on image size 768 for 5 epochs, and finally on image size 1024 for 5 more epochs. This achieves 0.80+ with a single EfficientNetB3. The intensity of ShiftScaleRotate basically doubled both in magnitude and probability during the process.</p>",
      "rawMarkdown": "Sorry, I don't feel very confident sharing the code right now as it is quite messy (I am learning PyTorch in these weeks) and also really similar to the [baseline](https://www.kaggle.com/code/surajsharan/ultramnist-starter-notebook). These are the main changes I've made to it:\n- Used [timm](https://www.kaggle.com/datasets/kozodoi/timm-pytorch-image-models) so I could have all models at hand\n- Removed resizing and added ShiftScaleRotate from Albumentations to the train transformations\n- Modified the training loop to keep track of train and validation metrics (accuracy and loss) so that I could plot their evolution in time later \n- Put the model in eval mode when predicting valid data\n- Added the possibility to save the class probabilities for the hold-out set and the test data after loading the best checkpoint\n\nI then resized and saved images in 384, 512, 768, and 1024 beforehand to speed things up. At each run of the notebook, I would load images of a certain size, change the intensity of augmentation and learning rate accordingly, then load the previous training checkpoint and resume from there. For example, first I trained on image size 384 for 15 epochs, then on image size 512 for 10 epochs, next on image size 768 for 5 epochs, and finally on image size 1024 for 5 more epochs. This achieves 0.80+ with a single EfficientNetB3. The intensity of ShiftScaleRotate basically doubled both in magnitude and probability during the process.",
      "votes": null
    },
    {
      "id": "1756748",
      "postDate": "04/15/2022 19:53:15",
      "content": "<p>Thanks for your interesting insights. They will complement my own first learnings of Pytorch. Many thanks also for the organizers who came up with the baseline Pytorch model to learn from. </p>\n<p>I stuck with 512px resolution and applied some RandomResizeCrop and Rotate augmentations to improve the baseline to finally get to 74% accuracy. On examining the erroneous train results I found that the model got the smallest digits wrong but didn't have the GPU quota to increase resolution to 1024px.</p>\n<p>Also interesting stuff from top 10 solutions. Will look up the Yolo object detection mechanics for sure.  </p>",
      "rawMarkdown": "Thanks for your interesting insights. They will complement my own first learnings of Pytorch. Many thanks also for the organizers who came up with the baseline Pytorch model to learn from. \n\nI stuck with 512px resolution and applied some RandomResizeCrop and Rotate augmentations to improve the baseline to finally get to 74% accuracy. On examining the erroneous train results I found that the model got the smallest digits wrong but didn't have the GPU quota to increase resolution to 1024px.\n\nAlso interesting stuff from top 10 solutions. Will look up the Yolo object detection mechanics for sure.",
      "votes": null
    },
    {
      "id": "1756752",
      "postDate": "04/15/2022 19:58:25",
      "content": "<p>Thanks for your interesting insights</p>",
      "rawMarkdown": "Thanks for your interesting insights",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1756448,
      "author_name": "devanshchowdhury",
      "author_url": "",
      "post_date": "04/15/2022 14:25:28",
      "content": "<p>Could you share the code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1756523,
          "author_name": "mviola",
          "author_url": "",
          "post_date": "04/15/2022 15:29:23",
          "content": "<p>Sorry, I don't feel very confident sharing the code right now as it is quite messy (I am learning PyTorch in these weeks) and also really similar to the <a href=\"https://www.kaggle.com/code/surajsharan/ultramnist-starter-notebook\" target=\"_blank\">baseline</a>. These are the main changes I've made to it:</p>\n<ul>\n<li>Used <a href=\"https://www.kaggle.com/datasets/kozodoi/timm-pytorch-image-models\" target=\"_blank\">timm</a> so I could have all models at hand</li>\n<li>Removed resizing and added ShiftScaleRotate from Albumentations to the train transformations</li>\n<li>Modified the training loop to keep track of train and validation metrics (accuracy and loss) so that I could plot their evolution in time later </li>\n<li>Put the model in eval mode when predicting valid data</li>\n<li>Added the possibility to save the class probabilities for the hold-out set and the test data after loading the best checkpoint</li>\n</ul>\n<p>I then resized and saved images in 384, 512, 768, and 1024 beforehand to speed things up. At each run of the notebook, I would load images of a certain size, change the intensity of augmentation and learning rate accordingly, then load the previous training checkpoint and resume from there. For example, first I trained on image size 384 for 15 epochs, then on image size 512 for 10 epochs, next on image size 768 for 5 epochs, and finally on image size 1024 for 5 more epochs. This achieves 0.80+ with a single EfficientNetB3. The intensity of ShiftScaleRotate basically doubled both in magnitude and probability during the process.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1756748,
      "author_name": "drumdroll",
      "author_url": "",
      "post_date": "04/15/2022 19:53:15",
      "content": "<p>Thanks for your interesting insights. They will complement my own first learnings of Pytorch. Many thanks also for the organizers who came up with the baseline Pytorch model to learn from. </p>\n<p>I stuck with 512px resolution and applied some RandomResizeCrop and Rotate augmentations to improve the baseline to finally get to 74% accuracy. On examining the erroneous train results I found that the model got the smallest digits wrong but didn't have the GPU quota to increase resolution to 1024px.</p>\n<p>Also interesting stuff from top 10 solutions. Will look up the Yolo object detection mechanics for sure.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1756752,
      "author_name": "vipin20",
      "author_url": "",
      "post_date": "04/15/2022 19:58:25",
      "content": "<p>Thanks for your interesting insights</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1756072": "Thanks to the organizers for creating a really interesting challenge inspired by a class of real-world problems and with great prizes. If I am being completely honest, I joined the competition with the old data just to have fun with OpenCV, but as it immediately became clear that it just solved the task and the dataset changed, I stuck around nonetheless as this was a nice opportunity to learn a bit of PyTorch as a loyal TensorFlow user. \n\n# Validation scheme\nMany reported that the test accuracy was lower than the validation one in the forum at the beginning, but after making sure train and test data came from the same distribution and submitting myself, I did not notice any issue. Since performing regular cross-validation was virtually impossible on large image sizes within the weekly GPU quota, I used a 10% hold-out set and it worked just fine.\n\n# Training methodology\nAfter thinking for some days about how to approach the innovation track, I did not come up with a good idea so I just trained a bunch of CNNs in PyTorch with Kaggle and Colab GPUs. To speed up training, I tried the progressive resizing as described in the [EfficientNetV2 paper](https://arxiv.org/pdf/2104.00298.pdf) and it worked just wonderfully. The dataset was noisy enough to prevent overfitting for a while, but in the end, I had to introduce some simple augmentation as ShiftScaleRotate, which was also increased during training alongside the image size.\n\n# Models\nMy solution is a pretty simple blend of my two top-scoring models, predicting the original dataset and the [inverted colour one](https://www.kaggle.com/competitions/ultra-mnist/discussion/313879). They are respectively an EfficientNetB3 and an EfficientNetV2B3 with image sizes 1024 for a total of four sets of predictions. The models all have CV/LB in the 0.8 range but are quite different. Using the hold-out dataset which contained the same image ids for all models, I found the best weights to combine these four predictions with Optuna and this gave me a decent boost, with my final submission scoring in the 0.85 range.\n\n# Key takeaways\n- CNNs are extremely powerful. Really, I would not have bet on such a high accuracy with such a noisy background and various digit sizes inside it, but I was totally wrong.\n- Pytorch is great and extremely beginner-friendly. I will continue to learn it for sure because its flexibility is just priceless.\n- Processing large image sizes is really time and memory consuming. This competition has brought to the fore situations where resizing is not an option if you want to achieve very high accuracy.\n\nThanks for taking the time to read this. I really hope to see an innovation track solution as this is an important problem, but I also don't mind seeing how the winners broke the competition again. Congratulations to all participants, let's keep learning!",
    "1756448": "Could you share the code?",
    "1756523": "Sorry, I don't feel very confident sharing the code right now as it is quite messy (I am learning PyTorch in these weeks) and also really similar to the [baseline](https://www.kaggle.com/code/surajsharan/ultramnist-starter-notebook). These are the main changes I've made to it:\n- Used [timm](https://www.kaggle.com/datasets/kozodoi/timm-pytorch-image-models) so I could have all models at hand\n- Removed resizing and added ShiftScaleRotate from Albumentations to the train transformations\n- Modified the training loop to keep track of train and validation metrics (accuracy and loss) so that I could plot their evolution in time later \n- Put the model in eval mode when predicting valid data\n- Added the possibility to save the class probabilities for the hold-out set and the test data after loading the best checkpoint\n\nI then resized and saved images in 384, 512, 768, and 1024 beforehand to speed things up. At each run of the notebook, I would load images of a certain size, change the intensity of augmentation and learning rate accordingly, then load the previous training checkpoint and resume from there. For example, first I trained on image size 384 for 15 epochs, then on image size 512 for 10 epochs, next on image size 768 for 5 epochs, and finally on image size 1024 for 5 more epochs. This achieves 0.80+ with a single EfficientNetB3. The intensity of ShiftScaleRotate basically doubled both in magnitude and probability during the process.",
    "1756748": "Thanks for your interesting insights. They will complement my own first learnings of Pytorch. Many thanks also for the organizers who came up with the baseline Pytorch model to learn from. \n\nI stuck with 512px resolution and applied some RandomResizeCrop and Rotate augmentations to improve the baseline to finally get to 74% accuracy. On examining the erroneous train results I found that the model got the smallest digits wrong but didn't have the GPU quota to increase resolution to 1024px.\n\nAlso interesting stuff from top 10 solutions. Will look up the Yolo object detection mechanics for sure.",
    "1756752": "Thanks for your interesting insights"
  },
  "source": "meta"
}