{
  "id": 168870,
  "title": "3rd place solution",
  "url": "/competitions/alaska2-image-steganalysis/writeups/kaizaburochubachi-3rd-place-solution",
  "author_name": "",
  "post_date": "2020-07-22T07:07:18.326313100Z",
  "votes": 45,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks for hosting this interesting, leakage free competition. I learned a lot from this competition. I also thank the other participants. I am very impressed with the unique solutions of the other participants. I will read all posted solutions.</p>\n\n<p>Here is a description of the my solution.</p>\n\n<h2>Validation</h2>\n\n<p>I split the file names randomly into 80% and 20% stratifying based on quality factor. Looking at the distribution of predictions for the test set, it seems that the percentage of cover images is relatively high. However, I thought that the AUC based metric is robust to such imbalance data, so I didn't take any special care of it at validation level.</p>\n\n<h2>Methods</h2>\n\n<p>I have created an ensemble of two models: one for inputting 3-channel color images and the other for feature engineering and inputting DCT coefficient.</p>\n\n<h3>3-channel color image model (validation score: 0.9382 (RGB), 0.9378 (YUV), 0.9364 (Lab))</h3>\n\n<ul>\n<li>I made 3 models that inputs RGB, YUV, and Lab, respectively.\n<ul><li>I used cv2 for loading and conversion. I also tried to use non-quantized YCrCb directly from DCT coefficient in the beginning, but the accuracy was almost the same as the YUV read with cv2, so I chose to use faster cv2.</li></ul></li>\n<li>4-class classification</li>\n<li>EfficientNet-b5\n<ul><li>I also tried RegNet, ResNeSt, HRNet, PyConv, and others, but EfficientNet was more accurate for this task than the other models which took a similar amount of time in my implementation.</li></ul></li>\n<li>Augmentation is Flip &amp; Rotate90 (8 types in total) and CutMix\n<ul><li>I tried to do CutMix with the data in different folder with the same file name, but the accuracy became worse. Did it become too difficult?</li></ul></li>\n<li>SGD + Cosine Annealing, 50 epoch per cycle\n<ul><li>Hyperparameters are almost identical to the EfficientNet experiment in <a href=\"https://arxiv.org/abs/1905.13214\">RegNet paper</a>.</li>\n<li>When I tried b0 and b2, I found that the accuracy of multiple cycles of Cosine Annealing was better than one cycle with the same total number of epochs. Especially, the gain at the second cycle was large. In the end, I ran 4 cycles of 50 epoch.</li></ul></li>\n</ul>\n\n<h3>DCT coefficient model (validation score: 0.9022)</h3>\n\n<p>I thought it would be difficult to make a differentiation in a model that uses RGB, so I attempted to create a network that directly handles the DCT coefficient by trial and error, but this was the only one that worked.</p>\n\n<ul>\n<li>Input features are\n<ul><li>one-hot encoding of the value {&lt;= -16, -15, -14, …, -2, -1, 1, 2, …, 14, 15, 16 &lt;=} </li>\n<li>positional encoding (I'm not sure if it helped)\n<ul><li>quantization matrix / 50</li>\n<li>a matrix such that matrix[i, j] = cos(pi * (i % 8) / 16) * cos(pi * (j % 8) / 16)</li></ul></li></ul></li>\n<li>Model architecture is \n<ul><li>EfficientNet-b2</li>\n<li>Change the first stride 2 to 1</li>\n<li>Change dilation of Conv2d to 8. This modification is applied Conv2d layers from the first (patched) stride 2 Conv2d to the next stride 2 Conv2d.</li></ul></li>\n<li>Augmentation is Flip &amp; Rotate90</li>\n<li>SGD + Cosine Annealing, 50 epoch, 1 cycle</li>\n<li>This model takes a long time to train, so I were not able to do a proper ablation study, but there was a difference of about 0.01 depending on whether the stride was changed or not.</li>\n</ul>\n\n<h3>Ensemble with MLP (validation score: 0.9399, private: 0.930)</h3>\n\n<p>Since outputs from the DCT coefficient model is quite different from outputs from 3-channel color image models, mixing it well improves the accuracy, but since it is considerably weaker than 3-channel models as a single model, it doesn't work well when simply averaging it. So I used MLP.</p>\n\n<ul>\n<li>Input the feature maps of DCT Binary, RGB, YUV, Lab after Average Pooling of each network. (There was a mistake, and softmax was applied to each feature maps.)</li>\n<li>To cancel the difference in scale between the networks, we use BatchNorm as follows.\n<code>\nx = torch.cat([bn(feat.reshape(-1, 1)).reshape(feat.shape) for feat, bn in zip(feats, self.bns)], dim=1)\n</code></li>\n</ul>\n\n<h3>Holdout stacking with LightGBM to push the last 0.001 (validation score: 0.9405, private: 0.931)</h3>\n\n<p>The holdout validation set has 15000 x 4 (Cover and Stegos) x 8 (flip / rotate90) = 4800000 examples, which is large enough to train LightGBM by splitting the holdout set into 5 folds. I used the MLP predictions and the t-SNE of each CNN feature map as features. The hyperparameter is determined by <a href=\"https://optuna.readthedocs.io/en/latest/reference/generated/optuna.integration.lightgbm.LightGBMTunerCV.html#optuna.integration.lightgbm.LightGBMTunerCV\">LightGBMTunerCV</a> further splitting the train data in each fold into 80% and 20%.</p>",
  "messages": [
    {
      "id": "939326",
      "postDate": "07/22/2020 07:07:18",
      "content": "<p>Thanks for hosting this interesting, leakage free competition. I learned a lot from this competition. I also thank the other participants. I am very impressed with the unique solutions of the other participants. I will read all posted solutions.</p>\n\n<p>Here is a description of the my solution.</p>\n\n<h2>Validation</h2>\n\n<p>I split the file names randomly into 80% and 20% stratifying based on quality factor. Looking at the distribution of predictions for the test set, it seems that the percentage of cover images is relatively high. However, I thought that the AUC based metric is robust to such imbalance data, so I didn't take any special care of it at validation level.</p>\n\n<h2>Methods</h2>\n\n<p>I have created an ensemble of two models: one for inputting 3-channel color images and the other for feature engineering and inputting DCT coefficient.</p>\n\n<h3>3-channel color image model (validation score: 0.9382 (RGB), 0.9378 (YUV), 0.9364 (Lab))</h3>\n\n<ul>\n<li>I made 3 models that inputs RGB, YUV, and Lab, respectively.\n<ul><li>I used cv2 for loading and conversion. I also tried to use non-quantized YCrCb directly from DCT coefficient in the beginning, but the accuracy was almost the same as the YUV read with cv2, so I chose to use faster cv2.</li></ul></li>\n<li>4-class classification</li>\n<li>EfficientNet-b5\n<ul><li>I also tried RegNet, ResNeSt, HRNet, PyConv, and others, but EfficientNet was more accurate for this task than the other models which took a similar amount of time in my implementation.</li></ul></li>\n<li>Augmentation is Flip &amp; Rotate90 (8 types in total) and CutMix\n<ul><li>I tried to do CutMix with the data in different folder with the same file name, but the accuracy became worse. Did it become too difficult?</li></ul></li>\n<li>SGD + Cosine Annealing, 50 epoch per cycle\n<ul><li>Hyperparameters are almost identical to the EfficientNet experiment in <a href=\"https://arxiv.org/abs/1905.13214\">RegNet paper</a>.</li>\n<li>When I tried b0 and b2, I found that the accuracy of multiple cycles of Cosine Annealing was better than one cycle with the same total number of epochs. Especially, the gain at the second cycle was large. In the end, I ran 4 cycles of 50 epoch.</li></ul></li>\n</ul>\n\n<h3>DCT coefficient model (validation score: 0.9022)</h3>\n\n<p>I thought it would be difficult to make a differentiation in a model that uses RGB, so I attempted to create a network that directly handles the DCT coefficient by trial and error, but this was the only one that worked.</p>\n\n<ul>\n<li>Input features are\n<ul><li>one-hot encoding of the value {&lt;= -16, -15, -14, …, -2, -1, 1, 2, …, 14, 15, 16 &lt;=} </li>\n<li>positional encoding (I'm not sure if it helped)\n<ul><li>quantization matrix / 50</li>\n<li>a matrix such that matrix[i, j] = cos(pi * (i % 8) / 16) * cos(pi * (j % 8) / 16)</li></ul></li></ul></li>\n<li>Model architecture is \n<ul><li>EfficientNet-b2</li>\n<li>Change the first stride 2 to 1</li>\n<li>Change dilation of Conv2d to 8. This modification is applied Conv2d layers from the first (patched) stride 2 Conv2d to the next stride 2 Conv2d.</li></ul></li>\n<li>Augmentation is Flip &amp; Rotate90</li>\n<li>SGD + Cosine Annealing, 50 epoch, 1 cycle</li>\n<li>This model takes a long time to train, so I were not able to do a proper ablation study, but there was a difference of about 0.01 depending on whether the stride was changed or not.</li>\n</ul>\n\n<h3>Ensemble with MLP (validation score: 0.9399, private: 0.930)</h3>\n\n<p>Since outputs from the DCT coefficient model is quite different from outputs from 3-channel color image models, mixing it well improves the accuracy, but since it is considerably weaker than 3-channel models as a single model, it doesn't work well when simply averaging it. So I used MLP.</p>\n\n<ul>\n<li>Input the feature maps of DCT Binary, RGB, YUV, Lab after Average Pooling of each network. (There was a mistake, and softmax was applied to each feature maps.)</li>\n<li>To cancel the difference in scale between the networks, we use BatchNorm as follows.\n<code>\nx = torch.cat([bn(feat.reshape(-1, 1)).reshape(feat.shape) for feat, bn in zip(feats, self.bns)], dim=1)\n</code></li>\n</ul>\n\n<h3>Holdout stacking with LightGBM to push the last 0.001 (validation score: 0.9405, private: 0.931)</h3>\n\n<p>The holdout validation set has 15000 x 4 (Cover and Stegos) x 8 (flip / rotate90) = 4800000 examples, which is large enough to train LightGBM by splitting the holdout set into 5 folds. I used the MLP predictions and the t-SNE of each CNN feature map as features. The hyperparameter is determined by <a href=\"https://optuna.readthedocs.io/en/latest/reference/generated/optuna.integration.lightgbm.LightGBMTunerCV.html#optuna.integration.lightgbm.LightGBMTunerCV\">LightGBMTunerCV</a> further splitting the train data in each fold into 80% and 20%.</p>",
      "rawMarkdown": "Thanks for hosting this interesting, leakage free competition. I learned a lot from this competition. I also thank the other participants. I am very impressed with the unique solutions of the other participants. I will read all posted solutions.\n\nHere is a description of the my solution.\n\n## Validation\n\nI split the file names randomly into 80% and 20% stratifying based on quality factor. Looking at the distribution of predictions for the test set, it seems that the percentage of cover images is relatively high. However, I thought that the AUC based metric is robust to such imbalance data, so I didn't take any special care of it at validation level.\n\n## Methods\n\nI have created an ensemble of two models: one for inputting 3-channel color images and the other for feature engineering and inputting DCT coefficient.\n\n### 3-channel color image model (validation score: 0.9382 (RGB), 0.9378 (YUV), 0.9364 (Lab))\n\n- I made 3 models that inputs RGB, YUV, and Lab, respectively.\n  - I used cv2 for loading and conversion. I also tried to use non-quantized YCrCb directly from DCT coefficient in the beginning, but the accuracy was almost the same as the YUV read with cv2, so I chose to use faster cv2.\n- 4-class classification\n- EfficientNet-b5\n  - I also tried RegNet, ResNeSt, HRNet, PyConv, and others, but EfficientNet was more accurate for this task than the other models which took a similar amount of time in my implementation.\n- Augmentation is Flip &amp; Rotate90 (8 types in total) and CutMix\n  - I tried to do CutMix with the data in different folder with the same file name, but the accuracy became worse. Did it become too difficult?\n- SGD + Cosine Annealing, 50 epoch per cycle\n  - Hyperparameters are almost identical to the EfficientNet experiment in [RegNet paper](https://arxiv.org/abs/1905.13214).\n  - When I tried b0 and b2, I found that the accuracy of multiple cycles of Cosine Annealing was better than one cycle with the same total number of epochs. Especially, the gain at the second cycle was large. In the end, I ran 4 cycles of 50 epoch.\n\n### DCT coefficient model (validation score: 0.9022)\n\nI thought it would be difficult to make a differentiation in a model that uses RGB, so I attempted to create a network that directly handles the DCT coefficient by trial and error, but this was the only one that worked.\n\n- Input features are\n  - one-hot encoding of the value {&lt;= -16, -15, -14, …, -2, -1, 1, 2, …, 14, 15, 16 &lt;=} \n  - positional encoding (I'm not sure if it helped)\n     - quantization matrix / 50\n     - a matrix such that matrix[i, j] = cos(pi * (i % 8) / 16) * cos(pi * (j % 8) / 16)\n- Model architecture is \n  - EfficientNet-b2\n  - Change the first stride 2 to 1\n  - Change dilation of Conv2d to 8. This modification is applied Conv2d layers from the first (patched) stride 2 Conv2d to the next stride 2 Conv2d.\n- Augmentation is Flip &amp; Rotate90\n- SGD + Cosine Annealing, 50 epoch, 1 cycle\n- This model takes a long time to train, so I were not able to do a proper ablation study, but there was a difference of about 0.01 depending on whether the stride was changed or not.\n\n### Ensemble with MLP (validation score: 0.9399, private: 0.930)\n\nSince outputs from the DCT coefficient model is quite different from outputs from 3-channel color image models, mixing it well improves the accuracy, but since it is considerably weaker than 3-channel models as a single model, it doesn't work well when simply averaging it. So I used MLP.\n\n- Input the feature maps of DCT Binary, RGB, YUV, Lab after Average Pooling of each network. (There was a mistake, and softmax was applied to each feature maps.)\n- To cancel the difference in scale between the networks, we use BatchNorm as follows.\n```\nx = torch.cat([bn(feat.reshape(-1, 1)).reshape(feat.shape) for feat, bn in zip(feats, self.bns)], dim=1)\n```\n\n### Holdout stacking with LightGBM to push the last 0.001 (validation score: 0.9405, private: 0.931)\n\nThe holdout validation set has 15000 x 4 (Cover and Stegos) x 8 (flip / rotate90) = 4800000 examples, which is large enough to train LightGBM by splitting the holdout set into 5 folds. I used the MLP predictions and the t-SNE of each CNN feature map as features. The hyperparameter is determined by [LightGBMTunerCV](https://optuna.readthedocs.io/en/latest/reference/generated/optuna.integration.lightgbm.LightGBMTunerCV.html#optuna.integration.lightgbm.LightGBMTunerCV) further splitting the train data in each fold into 80% and 20%.",
      "votes": null
    },
    {
      "id": "939397",
      "postDate": "07/22/2020 08:10:05",
      "content": "<p>Congrats <a href=\"/zaburo\">@zaburo</a> to you and Thanks for sharing the approach.</p>",
      "rawMarkdown": "Congrats @zaburo to you and Thanks for sharing the approach.",
      "votes": null
    },
    {
      "id": "940187",
      "postDate": "07/22/2020 18:13:20",
      "content": "<p>Nice solution! Thank you for sharing.</p>",
      "rawMarkdown": "Nice solution! Thank you for sharing.",
      "votes": null
    },
    {
      "id": "961282",
      "postDate": "08/07/2020 04:37:20",
      "content": "<p>We have published the code!\n<a href=\"https://github.com/pfnet-research/kaggle-alaska2-3rd-place-solution\">https://github.com/pfnet-research/kaggle-alaska2-3rd-place-solution</a></p>",
      "rawMarkdown": "We have published the code!\nhttps://github.com/pfnet-research/kaggle-alaska2-3rd-place-solution",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 939397,
      "author_name": "geekysaint",
      "author_url": "",
      "post_date": "07/22/2020 08:10:05",
      "content": "<p>Congrats <a href=\"/zaburo\">@zaburo</a> to you and Thanks for sharing the approach.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 940187,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "07/22/2020 18:13:20",
      "content": "<p>Nice solution! Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 961282,
      "author_name": "zaburo",
      "author_url": "",
      "post_date": "08/07/2020 04:37:20",
      "content": "<p>We have published the code!\n<a href=\"https://github.com/pfnet-research/kaggle-alaska2-3rd-place-solution\">https://github.com/pfnet-research/kaggle-alaska2-3rd-place-solution</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "939326": "Thanks for hosting this interesting, leakage free competition. I learned a lot from this competition. I also thank the other participants. I am very impressed with the unique solutions of the other participants. I will read all posted solutions.\n\nHere is a description of the my solution.\n\n## Validation\n\nI split the file names randomly into 80% and 20% stratifying based on quality factor. Looking at the distribution of predictions for the test set, it seems that the percentage of cover images is relatively high. However, I thought that the AUC based metric is robust to such imbalance data, so I didn't take any special care of it at validation level.\n\n## Methods\n\nI have created an ensemble of two models: one for inputting 3-channel color images and the other for feature engineering and inputting DCT coefficient.\n\n### 3-channel color image model (validation score: 0.9382 (RGB), 0.9378 (YUV), 0.9364 (Lab))\n\n- I made 3 models that inputs RGB, YUV, and Lab, respectively.\n  - I used cv2 for loading and conversion. I also tried to use non-quantized YCrCb directly from DCT coefficient in the beginning, but the accuracy was almost the same as the YUV read with cv2, so I chose to use faster cv2.\n- 4-class classification\n- EfficientNet-b5\n  - I also tried RegNet, ResNeSt, HRNet, PyConv, and others, but EfficientNet was more accurate for this task than the other models which took a similar amount of time in my implementation.\n- Augmentation is Flip &amp; Rotate90 (8 types in total) and CutMix\n  - I tried to do CutMix with the data in different folder with the same file name, but the accuracy became worse. Did it become too difficult?\n- SGD + Cosine Annealing, 50 epoch per cycle\n  - Hyperparameters are almost identical to the EfficientNet experiment in [RegNet paper](https://arxiv.org/abs/1905.13214).\n  - When I tried b0 and b2, I found that the accuracy of multiple cycles of Cosine Annealing was better than one cycle with the same total number of epochs. Especially, the gain at the second cycle was large. In the end, I ran 4 cycles of 50 epoch.\n\n### DCT coefficient model (validation score: 0.9022)\n\nI thought it would be difficult to make a differentiation in a model that uses RGB, so I attempted to create a network that directly handles the DCT coefficient by trial and error, but this was the only one that worked.\n\n- Input features are\n  - one-hot encoding of the value {&lt;= -16, -15, -14, …, -2, -1, 1, 2, …, 14, 15, 16 &lt;=} \n  - positional encoding (I'm not sure if it helped)\n     - quantization matrix / 50\n     - a matrix such that matrix[i, j] = cos(pi * (i % 8) / 16) * cos(pi * (j % 8) / 16)\n- Model architecture is \n  - EfficientNet-b2\n  - Change the first stride 2 to 1\n  - Change dilation of Conv2d to 8. This modification is applied Conv2d layers from the first (patched) stride 2 Conv2d to the next stride 2 Conv2d.\n- Augmentation is Flip &amp; Rotate90\n- SGD + Cosine Annealing, 50 epoch, 1 cycle\n- This model takes a long time to train, so I were not able to do a proper ablation study, but there was a difference of about 0.01 depending on whether the stride was changed or not.\n\n### Ensemble with MLP (validation score: 0.9399, private: 0.930)\n\nSince outputs from the DCT coefficient model is quite different from outputs from 3-channel color image models, mixing it well improves the accuracy, but since it is considerably weaker than 3-channel models as a single model, it doesn't work well when simply averaging it. So I used MLP.\n\n- Input the feature maps of DCT Binary, RGB, YUV, Lab after Average Pooling of each network. (There was a mistake, and softmax was applied to each feature maps.)\n- To cancel the difference in scale between the networks, we use BatchNorm as follows.\n```\nx = torch.cat([bn(feat.reshape(-1, 1)).reshape(feat.shape) for feat, bn in zip(feats, self.bns)], dim=1)\n```\n\n### Holdout stacking with LightGBM to push the last 0.001 (validation score: 0.9405, private: 0.931)\n\nThe holdout validation set has 15000 x 4 (Cover and Stegos) x 8 (flip / rotate90) = 4800000 examples, which is large enough to train LightGBM by splitting the holdout set into 5 folds. I used the MLP predictions and the t-SNE of each CNN feature map as features. The hyperparameter is determined by [LightGBMTunerCV](https://optuna.readthedocs.io/en/latest/reference/generated/optuna.integration.lightgbm.LightGBMTunerCV.html#optuna.integration.lightgbm.LightGBMTunerCV) further splitting the train data in each fold into 80% and 20%.",
    "939397": "Congrats @zaburo to you and Thanks for sharing the approach.",
    "940187": "Nice solution! Thank you for sharing.",
    "961282": "We have published the code!\nhttps://github.com/pfnet-research/kaggle-alaska2-3rd-place-solution"
  },
  "source": "meta"
}