{
  "id": 168537,
  "title": "[4th place] Definitely it's overfitting to the public LB. (0.948->0.930)",
  "url": "/competitions/alaska2-image-steganalysis/writeups/jonny-lee-4th-place-definitely-it-s-overfitting-to",
  "author_name": "",
  "post_date": "2020-07-22T12:17:28.683Z",
  "votes": 57,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Definitely it's overfitting to the public LB. But a solo gold is good enough to me. Maybe the best to me, because no need to submit my mess source code :) \nCongrats to everyone in this competition. We have learned a lot from it.</p>\n\n<p>Here is a brief of my solution.</p>\n\n<p>■ Using Tensorflow and TPU</p>\n\n<p>■ Preparation of the data\n Save DCT(512x512x16) as png into .tfrec file. (lossless)\n Save quality factor into .tfrec file.\n Save draft payload into .tfrec file. (If there any change in 8x8 area, count as 1)</p>\n\n<p>Transfer DCT into YCbCr in TF way at training time. By using this method, I can train one epoch in 15~20 mins for Effnet B0/B1, and 30~45 mins for B6/B7.\n The reason of using YCbCr is <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/150359#845167\">here</a>.</p>\n\n<p>■ Augmentation \n All of the augmentations are implemented in the TF way. (maybe there is something wrong here)</p>\n\n<p>Flip LR/UD\n Flip +/- (IMO it's available by using YCbCr)\n Multiply random number (0.98~1.02). (We are detecting the change of wavelets)\n Random shuffle the 24x24 / 32x32 / 40x40 blocks. IMO this can make model focus on wavelets of block but the contents of image. Keeping the boarder(16 pixels) because UERD always change this area.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2F86c070756b810c74293e6ada0707b630%2Fshuffle.png?generation=1595300985030345&amp;alt=media\" alt=\"\"></p>\n\n<p>No Rot90, because the quantization table is not diagonal symmetry. (maybe I was wrong)</p>\n\n<p>■ Models\n Changing one pixel by +/-1 in DCT will cause a change of 8x8 block in YCbCr, There are a lot of patterns to learn, so I tried B6/B7/B8 and the final result is the ensemble of them. The <a href=\"https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference\">best single model</a> is B6(Local/public/private: 0.940/0.940/0.929)</p>\n\n<p>Multiclass(4) + Aux loss (payload, mae)\n Quality factor as input. Not target.\n Using small learning rate at final stage.\n TTA: LR/UD/+-, 8 per image. (see Augmentation)</p>\n\n<p>■ Training\n I just focused on one fold and refined it on all data. Regrettably, it is overfitting to the public LB. By observing the public/private LB, I think many of others are like me except the winner. Congratulations again.</p>\n\n<p>If I trained more folds maybe I could ...... :) (It will take a long time and forget it now) </p>\n\n<p>■Update \n7/22  Add a link of <a href=\"https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference\">best single model</a>.</p>",
  "messages": [
    {
      "id": "937458",
      "postDate": "07/21/2020 03:24:01",
      "content": "<p>Definitely it's overfitting to the public LB. But a solo gold is good enough to me. Maybe the best to me, because no need to submit my mess source code :) \nCongrats to everyone in this competition. We have learned a lot from it.</p>\n\n<p>Here is a brief of my solution.</p>\n\n<p>■ Using Tensorflow and TPU</p>\n\n<p>■ Preparation of the data\n Save DCT(512x512x16) as png into .tfrec file. (lossless)\n Save quality factor into .tfrec file.\n Save draft payload into .tfrec file. (If there any change in 8x8 area, count as 1)</p>\n\n<p>Transfer DCT into YCbCr in TF way at training time. By using this method, I can train one epoch in 15~20 mins for Effnet B0/B1, and 30~45 mins for B6/B7.\n The reason of using YCbCr is <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/150359#845167\">here</a>.</p>\n\n<p>■ Augmentation \n All of the augmentations are implemented in the TF way. (maybe there is something wrong here)</p>\n\n<p>Flip LR/UD\n Flip +/- (IMO it's available by using YCbCr)\n Multiply random number (0.98~1.02). (We are detecting the change of wavelets)\n Random shuffle the 24x24 / 32x32 / 40x40 blocks. IMO this can make model focus on wavelets of block but the contents of image. Keeping the boarder(16 pixels) because UERD always change this area.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2F86c070756b810c74293e6ada0707b630%2Fshuffle.png?generation=1595300985030345&amp;alt=media\" alt=\"\"></p>\n\n<p>No Rot90, because the quantization table is not diagonal symmetry. (maybe I was wrong)</p>\n\n<p>■ Models\n Changing one pixel by +/-1 in DCT will cause a change of 8x8 block in YCbCr, There are a lot of patterns to learn, so I tried B6/B7/B8 and the final result is the ensemble of them. The <a href=\"https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference\">best single model</a> is B6(Local/public/private: 0.940/0.940/0.929)</p>\n\n<p>Multiclass(4) + Aux loss (payload, mae)\n Quality factor as input. Not target.\n Using small learning rate at final stage.\n TTA: LR/UD/+-, 8 per image. (see Augmentation)</p>\n\n<p>■ Training\n I just focused on one fold and refined it on all data. Regrettably, it is overfitting to the public LB. By observing the public/private LB, I think many of others are like me except the winner. Congratulations again.</p>\n\n<p>If I trained more folds maybe I could ...... :) (It will take a long time and forget it now) </p>\n\n<p>■Update \n7/22  Add a link of <a href=\"https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference\">best single model</a>.</p>",
      "rawMarkdown": "Definitely it's overfitting to the public LB. But a solo gold is good enough to me. Maybe the best to me, because no need to submit my mess source code :) \nCongrats to everyone in this competition. We have learned a lot from it.\n\nHere is a brief of my solution.\n\n■ Using Tensorflow and TPU\n\n■ Preparation of the data\n Save DCT(512x512x16) as png into .tfrec file. (lossless)\n Save quality factor into .tfrec file.\n Save draft payload into .tfrec file. (If there any change in 8x8 area, count as 1)\n\n Transfer DCT into YCbCr in TF way at training time. By using this method, I can train one epoch in 15~20 mins for Effnet B0/B1, and 30~45 mins for B6/B7.\n The reason of using YCbCr is [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/150359#845167).\n\n\n■ Augmentation \n All of the augmentations are implemented in the TF way. (maybe there is something wrong here)\n\n Flip LR/UD\n Flip +/- (IMO it's available by using YCbCr)\n Multiply random number (0.98~1.02). (We are detecting the change of wavelets)\n Random shuffle the 24x24 / 32x32 / 40x40 blocks. IMO this can make model focus on wavelets of block but the contents of image. Keeping the boarder(16 pixels) because UERD always change this area.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2F86c070756b810c74293e6ada0707b630%2Fshuffle.png?generation=1595300985030345&amp;alt=media)\n\n No Rot90, because the quantization table is not diagonal symmetry. (maybe I was wrong)\n\n\n■ Models\n Changing one pixel by +/-1 in DCT will cause a change of 8x8 block in YCbCr, There are a lot of patterns to learn, so I tried B6/B7/B8 and the final result is the ensemble of them. The [best single model](https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference) is B6(Local/public/private: 0.940/0.940/0.929)\n\n Multiclass(4) + Aux loss (payload, mae)\n Quality factor as input. Not target.\n Using small learning rate at final stage.\n TTA: LR/UD/+-, 8 per image. (see Augmentation)\n\n■ Training\n I just focused on one fold and refined it on all data. Regrettably, it is overfitting to the public LB. By observing the public/private LB, I think many of others are like me except the winner. Congratulations again.\n \n If I trained more folds maybe I could ...... :) (It will take a long time and forget it now) \n\n■Update \n7/22  Add a link of [best single model](https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference).",
      "votes": null
    },
    {
      "id": "937471",
      "postDate": "07/21/2020 03:38:12",
      "content": "<p>Thank you very much for sharing! I was wondering if you could publish your code. I also used TF/Keras on TPU, but I found the score much lower than same models with PyTorch on GPU. Also, I saw that you mentioned using efficientnet-b8. I tried to implement that with TF/Keras, but didn't work out for me. Thanks again for the great write up.</p>",
      "rawMarkdown": "Thank you very much for sharing! I was wondering if you could publish your code. I also used TF/Keras on TPU, but I found the score much lower than same models with PyTorch on GPU. Also, I saw that you mentioned using efficientnet-b8. I tried to implement that with TF/Keras, but didn't work out for me. Thanks again for the great write up.",
      "votes": null
    },
    {
      "id": "937508",
      "postDate": "07/21/2020 04:08:19",
      "content": "<p><a href=\"/wuliaokaola\">@wuliaokaola</a> </p>\n\n<p>\"No Rot90, because the quantization table is not diagonal symmetry. (maybe I was wrong)\"</p>\n\n<p>you can rotate the  quantization table too. you can refer to</p>\n\n<p><a href=\"https://linux.die.net/man/1/jpegtran\">https://linux.die.net/man/1/jpegtran</a>\n<a href=\"https://github.com/fhanau/Efficient-Compression-Tool/blob/8cc69fac9f26e3878298580222b890cc88d5e080/src/mozjpeg/transupp.c\">https://github.com/fhanau/Efficient-Compression-Tool/blob/8cc69fac9f26e3878298580222b890cc88d5e080/src/mozjpeg/transupp.c</a></p>",
      "rawMarkdown": "wuliaokaola \n\n\"No Rot90, because the quantization table is not diagonal symmetry. (maybe I was wrong)\"\n\nyou can rotate the  quantization table too. you can refer to\n\nhttps://linux.die.net/man/1/jpegtran\nhttps://github.com/fhanau/Efficient-Compression-Tool/blob/8cc69fac9f26e3878298580222b890cc88d5e080/src/mozjpeg/transupp.c",
      "votes": null
    },
    {
      "id": "937510",
      "postDate": "07/21/2020 04:09:34",
      "content": "<p>Sorry, I call it B8 but It's not real B8. I just change B7 more deeper and use larger dropout. </p>\n\n<p>efn.EfficientNet(\n                2.0, 3.8, 600, 0.55, 0.25,\n                input_tensor=x,\n                weights=None,\n                include_top=False\n            )</p>",
      "rawMarkdown": "Sorry, I call it B8 but It's not real B8. I just change B7 more deeper and use larger dropout. \n\nefn.EfficientNet(\n                2.0, 3.8, 600, 0.55, 0.25,\n                input_tensor=x,\n                weights=None,\n                include_top=False\n            )",
      "votes": null
    },
    {
      "id": "937523",
      "postDate": "07/21/2020 04:19:12",
      "content": "<p>I'm not sure it is a single rotation.</p>",
      "rawMarkdown": "I'm not sure it is a single rotation.",
      "votes": null
    },
    {
      "id": "937550",
      "postDate": "07/21/2020 04:36:31",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2Fb51702978100832b974499462e179c8d%2Fq.png?generation=1595306042130772&amp;alt=media\" alt=\"\">\nI know what you mean, but maybe the real stego is something different between vertical and horizon.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2Fb51702978100832b974499462e179c8d%2Fq.png?generation=1595306042130772&amp;alt=media)\nI know what you mean, but maybe the real stego is something different between vertical and horizon.",
      "votes": null
    },
    {
      "id": "937637",
      "postDate": "07/21/2020 05:39:30",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> </p>\n<p>Sorry for my naive question. But what are the wavelets ? <code>We are detecting the change of wavelets</code></p>",
      "rawMarkdown": "Congrats @wuliaokaola \n\nSorry for my naive question. But what are the wavelets ? `We are detecting the change of wavelets`",
      "votes": null
    },
    {
      "id": "937650",
      "postDate": "07/21/2020 05:43:19",
      "content": "<p>Thanks for sharing!\nI always wondered how you overfit.  👍 </p>",
      "rawMarkdown": "Thanks for sharing!\nI always wondered how you overfit.  👍",
      "votes": null
    },
    {
      "id": "937652",
      "postDate": "07/21/2020 05:43:51",
      "content": "<p>Congrats on gold medal and thanks for sharing your solution <a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a></p>",
      "rawMarkdown": "Congrats on gold medal and thanks for sharing your solution @wuliaokaola",
      "votes": null
    },
    {
      "id": "937695",
      "postDate": "07/21/2020 06:06:20",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "937716",
      "postDate": "07/21/2020 06:16:37",
      "content": "<p>I'm not familiar to it, too. But refer to this <a href=\"https://www.researchgate.net/publication/259639875_Universal_Distortion_Function_for_Steganography_in_an_Arbitrary_Domain\">paper</a>, UNIWARD make it difficult to be detected. And in this competition UNIWARD is the <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/155814\">most difficult one</a>.</p>",
      "rawMarkdown": "I'm not familiar to it, too. But refer to this [paper](https://www.researchgate.net/publication/259639875_Universal_Distortion_Function_for_Steganography_in_an_Arbitrary_Domain), UNIWARD make it difficult to be detected. And in this competition UNIWARD is the [most difficult one](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/155814).",
      "votes": null
    },
    {
      "id": "937804",
      "postDate": "07/21/2020 07:05:43",
      "content": "<p>k.π/2 rotations are fine as long as you rotate in the \"spatial\" domain, i.e. YCbCr or RGB. \nIf you want to rotate in the DCT domain, jpegtran has some implementation of \"lossless\" rotations, or you can refer to [http://www.ws.binghamton.edu/fridrich/Research/OneHot_Revised.pdf] Section VI-A  Data augmentation in DCT domain.</p>",
      "rawMarkdown": "k.π/2 rotations are fine as long as you rotate in the \"spatial\" domain, i.e. YCbCr or RGB. \nIf you want to rotate in the DCT domain, jpegtran has some implementation of \"lossless\" rotations, or you can refer to [http://www.ws.binghamton.edu/fridrich/Research/OneHot_Revised.pdf] Section VI-A  Data augmentation in DCT domain.",
      "votes": null
    },
    {
      "id": "937840",
      "postDate": "07/21/2020 07:30:31",
      "content": "<p><a href=\"/yousfi\">@yousfi</a> Thank you very much. But I still think the real stego is something different between vertical and horizon. In my experience, rot90 helps under 0.93(public LB), but makes the result worse above 0.93.</p>",
      "rawMarkdown": "yousfi Thank you very much. But I still think the real stego is something different between vertical and horizon. In my experience, rot90 helps under 0.93(public LB), but makes the result worse above 0.93.",
      "votes": null
    },
    {
      "id": "937847",
      "postDate": "07/21/2020 07:38:00",
      "content": "<p>I see what you mean, if the stego scheme has anisotropic cost assignments, then the stego image can be slightly different if rotated, but for the sake of augmentation, we can go away with that... For our models, D4 TTA always helped, as long as the net is trained with D4 augs.</p>",
      "rawMarkdown": "I see what you mean, if the stego scheme has anisotropic cost assignments, then the stego image can be slightly different if rotated, but for the sake of augmentation, we can go away with that... For our models, D4 TTA always helped, as long as the net is trained with D4 augs.",
      "votes": null
    },
    {
      "id": "937851",
      "postDate": "07/21/2020 07:44:06",
      "content": "<p>Thank you for your explanation. Maybe you are right.</p>",
      "rawMarkdown": "Thank you for your explanation. Maybe you are right.",
      "votes": null
    },
    {
      "id": "938161",
      "postDate": "07/21/2020 11:29:57",
      "content": "<p>Congrats! The method keeps the border and shuffles grid tiles was great. For UERD, the border information is important. I will try this type of shuffling to our models. Thanks for sharing.</p>",
      "rawMarkdown": "Congrats! The method keeps the border and shuffles grid tiles was great. For UERD, the border information is important. I will try this type of shuffling to our models. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "938300",
      "postDate": "07/21/2020 12:48:18",
      "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a><br>\nAlmost from the beginning to the end, you were the king on public LB. However, I'm curious to know 20 min per epoch, that's a lot of leverage to experiment. I think you were using kaggle TPU, it would be kind if you make a vanilla notebook of your work; not need to burn the TPU but simple and precise. Anyway, congratulation =)</p>",
      "rawMarkdown": "wuliaokaola\nAlmost from the beginning to the end, you were the king on public LB. However, I'm curious to know 20 min per epoch, that's a lot of leverage to experiment. I think you were using kaggle TPU, it would be kind if you make a vanilla notebook of your work; not need to burn the TPU but simple and precise. Anyway, congratulation =)",
      "votes": null
    },
    {
      "id": "938975",
      "postDate": "07/21/2020 23:09:23",
      "content": "<p>How did you incorporate the quality factor as input to efficient net? </p>",
      "rawMarkdown": "How did you incorporate the quality factor as input to efficient net?",
      "votes": null
    },
    {
      "id": "938995",
      "postDate": "07/22/2020 00:04:58",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    },
    {
      "id": "938999",
      "postDate": "07/22/2020 00:16:51",
      "content": "<p>Thank you.\n If you use .tfrec file,  20 min per epoch is easy to achieve. For TPU the bottleneck is I/O.</p>",
      "rawMarkdown": "Thank you.\n If you use .tfrec file,  20 min per epoch is easy to achieve. For TPU the bottleneck is I/O.",
      "votes": null
    },
    {
      "id": "939001",
      "postDate": "07/22/2020 00:20:48",
      "content": "<p>Treat quality factor as 4th channel. The input shape is 512x512x4.</p>",
      "rawMarkdown": "Treat quality factor as 4th channel. The input shape is 512x512x4.",
      "votes": null
    },
    {
      "id": "939012",
      "postDate": "07/22/2020 00:57:52",
      "content": "<p>Ah, I see thank you. Are you still able to use pretrained networks when you change the input channel dimension like this? </p>",
      "rawMarkdown": "Ah, I see thank you. Are you still able to use pretrained networks when you change the input channel dimension like this?",
      "votes": null
    },
    {
      "id": "939727",
      "postDate": "07/22/2020 12:13:44",
      "content": "<p><a href=\"/ipythonx\">@ipythonx</a> Here is an example of <a href=\"https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference\">inference</a>. If you want to use TPU, you should change GCS_DS_PATH_Test to \"KaggleDatasets()...\".</p>",
      "rawMarkdown": "ipythonx Here is an example of [inference](https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference). If you want to use TPU, you should change GCS_DS_PATH_Test to \"KaggleDatasets()...\".",
      "votes": null
    },
    {
      "id": "939742",
      "postDate": "07/22/2020 12:33:12",
      "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> <br>\nAwesome, thank you. =)</p>",
      "rawMarkdown": "wuliaokaola \nAwesome, thank you. =)",
      "votes": null
    },
    {
      "id": "941998",
      "postDate": "07/23/2020 14:35:56",
      "content": "<p><a href=\"/mr3543\">@mr3543</a> Yes. By using load_weights and set skip_mismatch to True.</p>",
      "rawMarkdown": "mr3543 Yes. By using load_weights and set skip_mismatch to True.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 937637,
      "author_name": "vishnurapps",
      "author_url": "",
      "post_date": "07/21/2020 05:39:30",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> </p>\n<p>Sorry for my naive question. But what are the wavelets ? <code>We are detecting the change of wavelets</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 937716,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/21/2020 06:16:37",
          "content": "<p>I'm not familiar to it, too. But refer to this <a href=\"https://www.researchgate.net/publication/259639875_Universal_Distortion_Function_for_Steganography_in_an_Arbitrary_Domain\">paper</a>, UNIWARD make it difficult to be detected. And in this competition UNIWARD is the <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/155814\">most difficult one</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937652,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "07/21/2020 05:43:51",
      "content": "<p>Congrats on gold medal and thanks for sharing your solution <a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 938300,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "07/21/2020 12:48:18",
      "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a><br>\nAlmost from the beginning to the end, you were the king on public LB. However, I'm curious to know 20 min per epoch, that's a lot of leverage to experiment. I think you were using kaggle TPU, it would be kind if you make a vanilla notebook of your work; not need to burn the TPU but simple and precise. Anyway, congratulation =)</p>",
      "votes": null,
      "replies": [
        {
          "id": 938999,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/22/2020 00:16:51",
          "content": "<p>Thank you.\n If you use .tfrec file,  20 min per epoch is easy to achieve. For TPU the bottleneck is I/O.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 939727,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/22/2020 12:13:44",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> Here is an example of <a href=\"https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference\">inference</a>. If you want to use TPU, you should change GCS_DS_PATH_Test to \"KaggleDatasets()...\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 939742,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/22/2020 12:33:12",
          "content": "<p><a href=\"https://www.kaggle.com/wuliaokaola\" target=\"_blank\">@wuliaokaola</a> <br>\nAwesome, thank you. =)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 938975,
      "author_name": "mr3543",
      "author_url": "",
      "post_date": "07/21/2020 23:09:23",
      "content": "<p>How did you incorporate the quality factor as input to efficient net? </p>",
      "votes": null,
      "replies": [
        {
          "id": 939001,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/22/2020 00:20:48",
          "content": "<p>Treat quality factor as 4th channel. The input shape is 512x512x4.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 939012,
          "author_name": "mr3543",
          "author_url": "",
          "post_date": "07/22/2020 00:57:52",
          "content": "<p>Ah, I see thank you. Are you still able to use pretrained networks when you change the input channel dimension like this? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 941998,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/23/2020 14:35:56",
          "content": "<p><a href=\"/mr3543\">@mr3543</a> Yes. By using load_weights and set skip_mismatch to True.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937471,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "07/21/2020 03:38:12",
      "content": "<p>Thank you very much for sharing! I was wondering if you could publish your code. I also used TF/Keras on TPU, but I found the score much lower than same models with PyTorch on GPU. Also, I saw that you mentioned using efficientnet-b8. I tried to implement that with TF/Keras, but didn't work out for me. Thanks again for the great write up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 937510,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/21/2020 04:09:34",
          "content": "<p>Sorry, I call it B8 but It's not real B8. I just change B7 more deeper and use larger dropout. </p>\n\n<p>efn.EfficientNet(\n                2.0, 3.8, 600, 0.55, 0.25,\n                input_tensor=x,\n                weights=None,\n                include_top=False\n            )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937695,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "07/21/2020 06:06:20",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937508,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/21/2020 04:08:19",
      "content": "<p><a href=\"/wuliaokaola\">@wuliaokaola</a> </p>\n\n<p>\"No Rot90, because the quantization table is not diagonal symmetry. (maybe I was wrong)\"</p>\n\n<p>you can rotate the  quantization table too. you can refer to</p>\n\n<p><a href=\"https://linux.die.net/man/1/jpegtran\">https://linux.die.net/man/1/jpegtran</a>\n<a href=\"https://github.com/fhanau/Efficient-Compression-Tool/blob/8cc69fac9f26e3878298580222b890cc88d5e080/src/mozjpeg/transupp.c\">https://github.com/fhanau/Efficient-Compression-Tool/blob/8cc69fac9f26e3878298580222b890cc88d5e080/src/mozjpeg/transupp.c</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 937523,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/21/2020 04:19:12",
          "content": "<p>I'm not sure it is a single rotation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937550,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/21/2020 04:36:31",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2Fb51702978100832b974499462e179c8d%2Fq.png?generation=1595306042130772&amp;alt=media\" alt=\"\">\nI know what you mean, but maybe the real stego is something different between vertical and horizon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937650,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "07/21/2020 05:43:19",
      "content": "<p>Thanks for sharing!\nI always wondered how you overfit.  👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937804,
      "author_name": "yousfi",
      "author_url": "",
      "post_date": "07/21/2020 07:05:43",
      "content": "<p>k.π/2 rotations are fine as long as you rotate in the \"spatial\" domain, i.e. YCbCr or RGB. \nIf you want to rotate in the DCT domain, jpegtran has some implementation of \"lossless\" rotations, or you can refer to [http://www.ws.binghamton.edu/fridrich/Research/OneHot_Revised.pdf] Section VI-A  Data augmentation in DCT domain.</p>",
      "votes": null,
      "replies": [
        {
          "id": 937840,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/21/2020 07:30:31",
          "content": "<p><a href=\"/yousfi\">@yousfi</a> Thank you very much. But I still think the real stego is something different between vertical and horizon. In my experience, rot90 helps under 0.93(public LB), but makes the result worse above 0.93.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937847,
          "author_name": "yousfi",
          "author_url": "",
          "post_date": "07/21/2020 07:38:00",
          "content": "<p>I see what you mean, if the stego scheme has anisotropic cost assignments, then the stego image can be slightly different if rotated, but for the sake of augmentation, we can go away with that... For our models, D4 TTA always helped, as long as the net is trained with D4 augs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937851,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/21/2020 07:44:06",
          "content": "<p>Thank you for your explanation. Maybe you are right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 938161,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "07/21/2020 11:29:57",
      "content": "<p>Congrats! The method keeps the border and shuffles grid tiles was great. For UERD, the border information is important. I will try this type of shuffling to our models. Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 938995,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "07/22/2020 00:04:58",
          "content": "<p>Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "937458": "Definitely it's overfitting to the public LB. But a solo gold is good enough to me. Maybe the best to me, because no need to submit my mess source code :) \nCongrats to everyone in this competition. We have learned a lot from it.\n\nHere is a brief of my solution.\n\n■ Using Tensorflow and TPU\n\n■ Preparation of the data\n Save DCT(512x512x16) as png into .tfrec file. (lossless)\n Save quality factor into .tfrec file.\n Save draft payload into .tfrec file. (If there any change in 8x8 area, count as 1)\n\n Transfer DCT into YCbCr in TF way at training time. By using this method, I can train one epoch in 15~20 mins for Effnet B0/B1, and 30~45 mins for B6/B7.\n The reason of using YCbCr is [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/150359#845167).\n\n\n■ Augmentation \n All of the augmentations are implemented in the TF way. (maybe there is something wrong here)\n\n Flip LR/UD\n Flip +/- (IMO it's available by using YCbCr)\n Multiply random number (0.98~1.02). (We are detecting the change of wavelets)\n Random shuffle the 24x24 / 32x32 / 40x40 blocks. IMO this can make model focus on wavelets of block but the contents of image. Keeping the boarder(16 pixels) because UERD always change this area.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2F86c070756b810c74293e6ada0707b630%2Fshuffle.png?generation=1595300985030345&amp;alt=media)\n\n No Rot90, because the quantization table is not diagonal symmetry. (maybe I was wrong)\n\n\n■ Models\n Changing one pixel by +/-1 in DCT will cause a change of 8x8 block in YCbCr, There are a lot of patterns to learn, so I tried B6/B7/B8 and the final result is the ensemble of them. The [best single model](https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference) is B6(Local/public/private: 0.940/0.940/0.929)\n\n Multiclass(4) + Aux loss (payload, mae)\n Quality factor as input. Not target.\n Using small learning rate at final stage.\n TTA: LR/UD/+-, 8 per image. (see Augmentation)\n\n■ Training\n I just focused on one fold and refined it on all data. Regrettably, it is overfitting to the public LB. By observing the public/private LB, I think many of others are like me except the winner. Congratulations again.\n \n If I trained more folds maybe I could ...... :) (It will take a long time and forget it now) \n\n■Update \n7/22  Add a link of [best single model](https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference).",
    "937471": "Thank you very much for sharing! I was wondering if you could publish your code. I also used TF/Keras on TPU, but I found the score much lower than same models with PyTorch on GPU. Also, I saw that you mentioned using efficientnet-b8. I tried to implement that with TF/Keras, but didn't work out for me. Thanks again for the great write up.",
    "937508": "wuliaokaola \n\n\"No Rot90, because the quantization table is not diagonal symmetry. (maybe I was wrong)\"\n\nyou can rotate the  quantization table too. you can refer to\n\nhttps://linux.die.net/man/1/jpegtran\nhttps://github.com/fhanau/Efficient-Compression-Tool/blob/8cc69fac9f26e3878298580222b890cc88d5e080/src/mozjpeg/transupp.c",
    "937510": "Sorry, I call it B8 but It's not real B8. I just change B7 more deeper and use larger dropout. \n\nefn.EfficientNet(\n                2.0, 3.8, 600, 0.55, 0.25,\n                input_tensor=x,\n                weights=None,\n                include_top=False\n            )",
    "937523": "I'm not sure it is a single rotation.",
    "937550": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2Fb51702978100832b974499462e179c8d%2Fq.png?generation=1595306042130772&amp;alt=media)\nI know what you mean, but maybe the real stego is something different between vertical and horizon.",
    "937637": "Congrats @wuliaokaola \n\nSorry for my naive question. But what are the wavelets ? `We are detecting the change of wavelets`",
    "937650": "Thanks for sharing!\nI always wondered how you overfit.  👍",
    "937652": "Congrats on gold medal and thanks for sharing your solution @wuliaokaola",
    "937695": "Thank you!",
    "937716": "I'm not familiar to it, too. But refer to this [paper](https://www.researchgate.net/publication/259639875_Universal_Distortion_Function_for_Steganography_in_an_Arbitrary_Domain), UNIWARD make it difficult to be detected. And in this competition UNIWARD is the [most difficult one](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/155814).",
    "937804": "k.π/2 rotations are fine as long as you rotate in the \"spatial\" domain, i.e. YCbCr or RGB. \nIf you want to rotate in the DCT domain, jpegtran has some implementation of \"lossless\" rotations, or you can refer to [http://www.ws.binghamton.edu/fridrich/Research/OneHot_Revised.pdf] Section VI-A  Data augmentation in DCT domain.",
    "937840": "yousfi Thank you very much. But I still think the real stego is something different between vertical and horizon. In my experience, rot90 helps under 0.93(public LB), but makes the result worse above 0.93.",
    "937847": "I see what you mean, if the stego scheme has anisotropic cost assignments, then the stego image can be slightly different if rotated, but for the sake of augmentation, we can go away with that... For our models, D4 TTA always helped, as long as the net is trained with D4 augs.",
    "937851": "Thank you for your explanation. Maybe you are right.",
    "938161": "Congrats! The method keeps the border and shuffles grid tiles was great. For UERD, the border information is important. I will try this type of shuffling to our models. Thanks for sharing.",
    "938300": "wuliaokaola\nAlmost from the beginning to the end, you were the king on public LB. However, I'm curious to know 20 min per epoch, that's a lot of leverage to experiment. I think you were using kaggle TPU, it would be kind if you make a vanilla notebook of your work; not need to burn the TPU but simple and precise. Anyway, congratulation =)",
    "938975": "How did you incorporate the quality factor as input to efficient net?",
    "938995": "Thank you.",
    "938999": "Thank you.\n If you use .tfrec file,  20 min per epoch is easy to achieve. For TPU the bottleneck is I/O.",
    "939001": "Treat quality factor as 4th channel. The input shape is 512x512x4.",
    "939012": "Ah, I see thank you. Are you still able to use pretrained networks when you change the input channel dimension like this?",
    "939727": "ipythonx Here is an example of [inference](https://www.kaggle.com/wuliaokaola/alaska2-best-b6-inference). If you want to use TPU, you should change GCS_DS_PATH_Test to \"KaggleDatasets()...\".",
    "939742": "wuliaokaola \nAwesome, thank you. =)",
    "941998": "mr3543 Yes. By using load_weights and set skip_mismatch to True."
  },
  "source": "meta"
}