{
  "id": 168519,
  "title": "[8th Place] Brief solution.",
  "url": "/competitions/alaska2-image-steganalysis/writeups/bandabi-8th-place-brief-solution",
  "author_name": "",
  "post_date": "2021-11-05T17:18:13.223Z",
  "votes": 37,
  "comment_count": 21,
  "views": 0,
  "content": "<h1>Before posting</h1>\n<p>Before posting, Thanks for hosting this great competition. And also, Thanks to our great teammates: <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, <a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a>, <a href=\"https://www.kaggle.com/wonhyukahn\" target=\"_blank\">@wonhyukahn</a>, <a href=\"https://www.kaggle.com/bhjang\" target=\"_blank\">@bhjang</a>.</p>\n<p>Sorry for this brief solution, but details will be updated soon.</p>\n<h1>Pipeline</h1>\n<p>Our pipelines are below:</p>\n<blockquote>\n  <p>EDA -&gt; CV strategy -&gt; Data -&gt; Augmentations -&gt; Models -&gt; Ensemble</p>\n</blockquote>\n<h1>1. EDA</h1>\n<blockquote>\n  <p>will be updated soon</p>\n</blockquote>\n<h1>2. CV strategy</h1>\n<blockquote>\n  <ul>\n  <li>Stratified KFold based on quality factor and stego information insertion amount</li>\n  </ul>\n</blockquote>\n<p>Information insertion amount could be calculated below:</p>\n<blockquote>\n  <p>a = difference between cover and stego to DCT coefficients per channel<br>\n  b = number of non-zero DCT coefficients per channel<br>\n  information insertion amount = a / b</p>\n</blockquote>\n<p>Then, After binning the information into 10, it was divided into k folds for each quality factor. At early stage we used 5 folds, but later 10 folds.</p>\n<p>Because it took so long to train model, we didn't train all folds. One of the k folds was selected and used as hold-out.</p>\n<h1>3. Data</h1>\n<blockquote>\n  <ul>\n  <li>Image: From RGB to ReconstructedRGB</li>\n  <li>Label: From binary to multi-label</li>\n  </ul>\n</blockquote>\n<p>For image, we use reconstructed RGB instead of RGB. When jpeg image is loaded using cv2, the small value is rounded and truncated. Since Stego information is a very fine information, it was important to preserve these values. Thus, We create an image without rounding and truncation using jpegio library.</p>\n<p>For label, as mentioned earlier in discussions, multi-label classification was used rather than binary label. we used 10 targets (Cover, 3 quality factors * 3 stego algorithms).</p>\n<h1>4. Augmentations</h1>\n<blockquote>\n  <ul>\n  <li>Flip, Rotate90, Cutout, GridShuffle, GridDropout</li>\n  </ul>\n</blockquote>\n<p>I tried a lot for augmentation experiments.<br>\nFirst most importantly, we remove augmentations that changes the original pixel value. Thus, we use <strong>Flip, Rotate90, Cutout, GridShuffle, and GridDropout</strong>. </p>\n<p>The best result was GridShuffle. This augmentation came from thinking about the answer to the next question: How can our model focus on fine stego information, not the shape of the object in image? <br>\nTo answer these question, We tried shuffle pixels then shuffle 8x8 tiles. In JPEG image, 8x8 tiles are important because DCT information is included in 8x8 tiles. Stego information is inserted at the DCT level, and since the DCT information is continuous within 8x8 tiles, We decided to mix these 8x8 tiles. <br>\nIf the unit size of tile is set to 8x8, a total of 64*64 tiles are generated. Mixing all these tiles breaks the <strong>texture information</strong> along with the shape information, which hinders the model from being trained. Thus, we chose a suitable sized tile that is not too small and trained the model.<br>\nThe GridDropout is similar.</p>\n<p>code is below:</p>\n<ul>\n<li>GridShuffle(Albumentations)</li>\n</ul>\n<blockquote>\n  <p>grid_shuffle = OneOf([<br>\n          RandomGridShuffle((2, 2), p=1.0),<br>\n          RandomGridShuffle((2, 4), p=1.0),<br>\n          RandomGridShuffle((2, 8), p=1.0),<br>\n          RandomGridShuffle((2, 16), p=1.0),<br>\n          RandomGridShuffle((2, 32), p=1.0),<br>\n          RandomGridShuffle((4, 4), p=1.0),<br>\n          RandomGridShuffle((4, 8), p=1.0),<br>\n          RandomGridShuffle((4, 16), p=1.0),<br>\n          RandomGridShuffle((4, 32), p=1.0),<br>\n          RandomGridShuffle((8, 8), p=1.0),<br>\n          RandomGridShuffle((8, 16), p=1.0),<br>\n          RandomGridShuffle((8, 32), p=1.0),<br>\n          RandomGridShuffle((16, 16), p=1.0),<br>\n          RandomGridShuffle((16, 32), p=1.0),<br>\n          RandomGridShuffle((32, 32), p=1.0),<br>\n      ], p=0.7)</p>\n</blockquote>\n<h1>5. Models</h1>\n<blockquote>\n  <ul>\n  <li>Framework: Pytorch</li>\n  <li>Backbone: Efficientnet b0, b0(double resolution), b3, b4, b7, resnext</li>\n  <li>Scheduler: ReduceOnplateau</li>\n  <li>Loss: CrossEntropy Loss</li>\n  <li>Optimizer: AdamP, AdamW</li>\n  <li>Epochs: 100+</li>\n  <li>Use TTA 8x</li>\n  </ul>\n</blockquote>\n<h1>6. ETC</h1>\n<blockquote>\n  <ul>\n  <li>Freeze backbone and finetune fc layer</li>\n  <li>Use all data as training</li>\n  <li>Flipped Stego Images</li>\n  </ul>\n</blockquote>\n<h1>7. Scores</h1>\n<ul>\n<li>Best Single Model in public leaderboard - Public LB: 945, Private LB: 925</li>\n<li>Best Single Model in private leaderboard - Public LB: 940, Private LB: 928</li>\n<li>Ensemble - Public LB: 944, Private LB: 929</li>\n</ul>",
  "messages": [
    {
      "id": "937386",
      "postDate": "07/21/2020 01:30:38",
      "content": "<h1>Before posting</h1>\n<p>Before posting, Thanks for hosting this great competition. And also, Thanks to our great teammates: <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a>, <a href=\"https://www.kaggle.com/youhanlee\" target=\"_blank\">@youhanlee</a>, <a href=\"https://www.kaggle.com/wonhyukahn\" target=\"_blank\">@wonhyukahn</a>, <a href=\"https://www.kaggle.com/bhjang\" target=\"_blank\">@bhjang</a>.</p>\n<p>Sorry for this brief solution, but details will be updated soon.</p>\n<h1>Pipeline</h1>\n<p>Our pipelines are below:</p>\n<blockquote>\n  <p>EDA -&gt; CV strategy -&gt; Data -&gt; Augmentations -&gt; Models -&gt; Ensemble</p>\n</blockquote>\n<h1>1. EDA</h1>\n<blockquote>\n  <p>will be updated soon</p>\n</blockquote>\n<h1>2. CV strategy</h1>\n<blockquote>\n  <ul>\n  <li>Stratified KFold based on quality factor and stego information insertion amount</li>\n  </ul>\n</blockquote>\n<p>Information insertion amount could be calculated below:</p>\n<blockquote>\n  <p>a = difference between cover and stego to DCT coefficients per channel<br>\n  b = number of non-zero DCT coefficients per channel<br>\n  information insertion amount = a / b</p>\n</blockquote>\n<p>Then, After binning the information into 10, it was divided into k folds for each quality factor. At early stage we used 5 folds, but later 10 folds.</p>\n<p>Because it took so long to train model, we didn't train all folds. One of the k folds was selected and used as hold-out.</p>\n<h1>3. Data</h1>\n<blockquote>\n  <ul>\n  <li>Image: From RGB to ReconstructedRGB</li>\n  <li>Label: From binary to multi-label</li>\n  </ul>\n</blockquote>\n<p>For image, we use reconstructed RGB instead of RGB. When jpeg image is loaded using cv2, the small value is rounded and truncated. Since Stego information is a very fine information, it was important to preserve these values. Thus, We create an image without rounding and truncation using jpegio library.</p>\n<p>For label, as mentioned earlier in discussions, multi-label classification was used rather than binary label. we used 10 targets (Cover, 3 quality factors * 3 stego algorithms).</p>\n<h1>4. Augmentations</h1>\n<blockquote>\n  <ul>\n  <li>Flip, Rotate90, Cutout, GridShuffle, GridDropout</li>\n  </ul>\n</blockquote>\n<p>I tried a lot for augmentation experiments.<br>\nFirst most importantly, we remove augmentations that changes the original pixel value. Thus, we use <strong>Flip, Rotate90, Cutout, GridShuffle, and GridDropout</strong>. </p>\n<p>The best result was GridShuffle. This augmentation came from thinking about the answer to the next question: How can our model focus on fine stego information, not the shape of the object in image? <br>\nTo answer these question, We tried shuffle pixels then shuffle 8x8 tiles. In JPEG image, 8x8 tiles are important because DCT information is included in 8x8 tiles. Stego information is inserted at the DCT level, and since the DCT information is continuous within 8x8 tiles, We decided to mix these 8x8 tiles. <br>\nIf the unit size of tile is set to 8x8, a total of 64*64 tiles are generated. Mixing all these tiles breaks the <strong>texture information</strong> along with the shape information, which hinders the model from being trained. Thus, we chose a suitable sized tile that is not too small and trained the model.<br>\nThe GridDropout is similar.</p>\n<p>code is below:</p>\n<ul>\n<li>GridShuffle(Albumentations)</li>\n</ul>\n<blockquote>\n  <p>grid_shuffle = OneOf([<br>\n          RandomGridShuffle((2, 2), p=1.0),<br>\n          RandomGridShuffle((2, 4), p=1.0),<br>\n          RandomGridShuffle((2, 8), p=1.0),<br>\n          RandomGridShuffle((2, 16), p=1.0),<br>\n          RandomGridShuffle((2, 32), p=1.0),<br>\n          RandomGridShuffle((4, 4), p=1.0),<br>\n          RandomGridShuffle((4, 8), p=1.0),<br>\n          RandomGridShuffle((4, 16), p=1.0),<br>\n          RandomGridShuffle((4, 32), p=1.0),<br>\n          RandomGridShuffle((8, 8), p=1.0),<br>\n          RandomGridShuffle((8, 16), p=1.0),<br>\n          RandomGridShuffle((8, 32), p=1.0),<br>\n          RandomGridShuffle((16, 16), p=1.0),<br>\n          RandomGridShuffle((16, 32), p=1.0),<br>\n          RandomGridShuffle((32, 32), p=1.0),<br>\n      ], p=0.7)</p>\n</blockquote>\n<h1>5. Models</h1>\n<blockquote>\n  <ul>\n  <li>Framework: Pytorch</li>\n  <li>Backbone: Efficientnet b0, b0(double resolution), b3, b4, b7, resnext</li>\n  <li>Scheduler: ReduceOnplateau</li>\n  <li>Loss: CrossEntropy Loss</li>\n  <li>Optimizer: AdamP, AdamW</li>\n  <li>Epochs: 100+</li>\n  <li>Use TTA 8x</li>\n  </ul>\n</blockquote>\n<h1>6. ETC</h1>\n<blockquote>\n  <ul>\n  <li>Freeze backbone and finetune fc layer</li>\n  <li>Use all data as training</li>\n  <li>Flipped Stego Images</li>\n  </ul>\n</blockquote>\n<h1>7. Scores</h1>\n<ul>\n<li>Best Single Model in public leaderboard - Public LB: 945, Private LB: 925</li>\n<li>Best Single Model in private leaderboard - Public LB: 940, Private LB: 928</li>\n<li>Ensemble - Public LB: 944, Private LB: 929</li>\n</ul>",
      "rawMarkdown": "# Before posting\nBefore posting, Thanks for hosting this great competition. And also, Thanks to our great teammates: @titericz, @youhanlee, @wonhyukahn, @bhjang.\n\nSorry for this brief solution, but details will be updated soon.\n\n# Pipeline\nOur pipelines are below:\n&gt; EDA -&gt; CV strategy -&gt; Data -&gt; Augmentations -&gt; Models -&gt; Ensemble\n\n# 1. EDA\n&gt; will be updated soon\n\n# 2. CV strategy\n&gt; - Stratified KFold based on quality factor and stego information insertion amount\n\nInformation insertion amount could be calculated below:\n&gt; a = difference between cover and stego to DCT coefficients per channel\nb = number of non-zero DCT coefficients per channel\ninformation insertion amount = a / b\n\nThen, After binning the information into 10, it was divided into k folds for each quality factor. At early stage we used 5 folds, but later 10 folds.\n\nBecause it took so long to train model, we didn't train all folds. One of the k folds was selected and used as hold-out.\n\n# 3. Data\n&gt; - Image: From RGB to ReconstructedRGB\n- Label: From binary to multi-label\n\nFor image, we use reconstructed RGB instead of RGB. When jpeg image is loaded using cv2, the small value is rounded and truncated. Since Stego information is a very fine information, it was important to preserve these values. Thus, We create an image without rounding and truncation using jpegio library.\n\nFor label, as mentioned earlier in discussions, multi-label classification was used rather than binary label. we used 10 targets (Cover, 3 quality factors * 3 stego algorithms).\n\n# 4. Augmentations\n&gt; - Flip, Rotate90, Cutout, GridShuffle, GridDropout\n\nI tried a lot for augmentation experiments.\nFirst most importantly, we remove augmentations that changes the original pixel value. Thus, we use **Flip, Rotate90, Cutout, GridShuffle, and GridDropout**. \n\nThe best result was GridShuffle. This augmentation came from thinking about the answer to the next question: How can our model focus on fine stego information, not the shape of the object in image? \nTo answer these question, We tried shuffle pixels then shuffle 8x8 tiles. In JPEG image, 8x8 tiles are important because DCT information is included in 8x8 tiles. Stego information is inserted at the DCT level, and since the DCT information is continuous within 8x8 tiles, We decided to mix these 8x8 tiles. \nIf the unit size of tile is set to 8x8, a total of 64*64 tiles are generated. Mixing all these tiles breaks the **texture information** along with the shape information, which hinders the model from being trained. Thus, we chose a suitable sized tile that is not too small and trained the model.\nThe GridDropout is similar.\n\ncode is below:\n\n* GridShuffle(Albumentations)\n&gt; grid_shuffle = OneOf([\n        RandomGridShuffle((2, 2), p=1.0),\n        RandomGridShuffle((2, 4), p=1.0),\n        RandomGridShuffle((2, 8), p=1.0),\n        RandomGridShuffle((2, 16), p=1.0),\n        RandomGridShuffle((2, 32), p=1.0),\n        RandomGridShuffle((4, 4), p=1.0),\n        RandomGridShuffle((4, 8), p=1.0),\n        RandomGridShuffle((4, 16), p=1.0),\n        RandomGridShuffle((4, 32), p=1.0),\n        RandomGridShuffle((8, 8), p=1.0),\n        RandomGridShuffle((8, 16), p=1.0),\n        RandomGridShuffle((8, 32), p=1.0),\n        RandomGridShuffle((16, 16), p=1.0),\n        RandomGridShuffle((16, 32), p=1.0),\n        RandomGridShuffle((32, 32), p=1.0),\n    ], p=0.7)\n\n\n# 5. Models\n&gt; - Framework: Pytorch\n- Backbone: Efficientnet b0, b0(double resolution), b3, b4, b7, resnext\n- Scheduler: ReduceOnplateau\n- Loss: CrossEntropy Loss\n- Optimizer: AdamP, AdamW\n- Epochs: 100+\n- Use TTA 8x\n\n# 6. ETC\n&gt; - Freeze backbone and finetune fc layer\n- Use all data as training\n- Flipped Stego Images\n\n# 7. Scores\n* Best Single Model in public leaderboard - Public LB: 945, Private LB: 925\n* Best Single Model in private leaderboard - Public LB: 940, Private LB: 928\n* Ensemble - Public LB: 944, Private LB: 929",
      "votes": null
    },
    {
      "id": "937411",
      "postDate": "07/21/2020 02:12:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a> ,</p>\n<p>Wonderful work. What I found most difficult was training took a lot of time especially B7 with pytorch. Can you help me with insight over your approach of tackling it. Was keras better than pytorch ? Did you use colab (GPU/TPU) to train ?</p>",
      "rawMarkdown": "Hi @songwonho ,\n\nWonderful work. What I found most difficult was training took a lot of time especially B7 with pytorch. Can you help me with insight over your approach of tackling it. Was keras better than pytorch ? Did you use colab (GPU/TPU) to train ?",
      "votes": null
    },
    {
      "id": "937469",
      "postDate": "07/21/2020 03:37:50",
      "content": "<p>thanks, We use pytorch. I think that large batch and large model would be great for this competition, so i tried to use TPU. However the competition was almost over, so I couldn't use it eventually with Reconstructed RGB.</p>",
      "rawMarkdown": "thanks, We use pytorch. I think that large batch and large model would be great for this competition, so i tried to use TPU. However the competition was almost over, so I couldn't use it eventually with Reconstructed RGB.",
      "votes": null
    },
    {
      "id": "937513",
      "postDate": "07/21/2020 04:10:06",
      "content": "<p>Thanks for a brief of the solution so early, I am very eager to learn although cv competitions are really computationally heavy and I don't have any gpus, can you give any advice/suggestions etc regarding this to a newbie like me\nThanks </p>",
      "rawMarkdown": "Thanks for a brief of the solution so early, I am very eager to learn although cv competitions are really computationally heavy and I don't have any gpus, can you give any advice/suggestions etc regarding this to a newbie like me\nThanks",
      "votes": null
    },
    {
      "id": "937560",
      "postDate": "07/21/2020 04:43:00",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a>  for the rank and thanks to share your pipeline but i want to ask something that why you have froze the backbone and fc layers ??</p>",
      "rawMarkdown": "congrats @songwonho  for the rank and thanks to share your pipeline but i want to ask something that why you have froze the backbone and fc layers ??",
      "votes": null
    },
    {
      "id": "937574",
      "postDate": "07/21/2020 04:54:26",
      "content": "<p>Thanks. It`s a genius from giba. When we freeze backbone, we can train a model with a large batch size. Also, stego information is very small, we finetune the fully connected layer more through small learning rate.</p>",
      "rawMarkdown": "Thanks. It`s a genius from giba. When we freeze backbone, we can train a model with a large batch size. Also, stego information is very small, we finetune the fully connected layer more through small learning rate.",
      "votes": null
    },
    {
      "id": "937589",
      "postDate": "07/21/2020 05:05:26",
      "content": "<p>Thanks. Nowadays cv competitions require huge resources. I think the fastest way is to buy a new powerful gpu, but it`s not easy. Here are some of method that have helped me.</p>\n\n<ol>\n<li><p>Use TPU\nWe can take advantage of a good TPU machine through Kaggle or Colab.</p></li>\n<li><p>Fast Experiment Settings\nIn the early stage, I set up an experimental environment that can see results quickly. For example, i tried a lot of data augmentations using efficientnet-b0 and large learning rate for fast convergence. Record good experiment results at this stage, and later apply good experiments through resources such as GCP or colab. What is important here is that the results in the fast experimental environment should align the results in the real experiment.</p></li>\n<li><p>Team-up\nIf we have a good idea, we can do it with colleagues with good equipment.</p></li>\n</ol>",
      "rawMarkdown": "Thanks. Nowadays cv competitions require huge resources. I think the fastest way is to buy a new powerful gpu, but it`s not easy. Here are some of method that have helped me.\n\n1. Use TPU\nWe can take advantage of a good TPU machine through Kaggle or Colab.\n\n2. Fast Experiment Settings\nIn the early stage, I set up an experimental environment that can see results quickly. For example, i tried a lot of data augmentations using efficientnet-b0 and large learning rate for fast convergence. Record good experiment results at this stage, and later apply good experiments through resources such as GCP or colab. What is important here is that the results in the fast experimental environment should align the results in the real experiment.\n\n3. Team-up\nIf we have a good idea, we can do it with colleagues with good equipment.",
      "votes": null
    },
    {
      "id": "937618",
      "postDate": "07/21/2020 05:29:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a> and team.</p>\n<p>I have a doubt regarding how grid shuffle and dropout help. When I checked using dct most of the information were in one part of the image and doing the shuffle disrupts the pattern. Could you please help me understand how it helped ?</p>",
      "rawMarkdown": "Congrats @songwonho and team.\n\nI have a doubt regarding how grid shuffle and dropout help. When I checked using dct most of the information were in one part of the image and doing the shuffle disrupts the pattern. Could you please help me understand how it helped ?",
      "votes": null
    },
    {
      "id": "937621",
      "postDate": "07/21/2020 05:29:52",
      "content": "<p>Thank you for the insights. Its really helpful.</p>",
      "rawMarkdown": "Thank you for the insights. Its really helpful.",
      "votes": null
    },
    {
      "id": "937636",
      "postDate": "07/21/2020 05:39:16",
      "content": "<p>8x8 tiles was import in dct field. stego information was spread in 8x8 tiles and has a continuous distribution. Thus, while preserving this information, I proceeded with grid shuffle and grid dropout. My intention was that model could see stego information instead of object shape in image using 8x8 tile grid related augmentations. Details will be added to text soon!</p>",
      "rawMarkdown": "8x8 tiles was import in dct field. stego information was spread in 8x8 tiles and has a continuous distribution. Thus, while preserving this information, I proceeded with grid shuffle and grid dropout. My intention was that model could see stego information instead of object shape in image using 8x8 tile grid related augmentations. Details will be added to text soon!",
      "votes": null
    },
    {
      "id": "937639",
      "postDate": "07/21/2020 05:39:56",
      "content": "<p>Thanks for sharing!\n😊 </p>",
      "rawMarkdown": "Thanks for sharing!\n😊",
      "votes": null
    },
    {
      "id": "937649",
      "postDate": "07/21/2020 05:43:06",
      "content": "<p>thanks.</p>",
      "rawMarkdown": "thanks.",
      "votes": null
    },
    {
      "id": "937659",
      "postDate": "07/21/2020 05:45:42",
      "content": "<p>Congrats to you and team on 8th place and thanks for sharing your solution <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a>!</p>",
      "rawMarkdown": "Congrats to you and team on 8th place and thanks for sharing your solution @songwonho!",
      "votes": null
    },
    {
      "id": "937664",
      "postDate": "07/21/2020 05:47:15",
      "content": "<p>Thanks !</p>",
      "rawMarkdown": "Thanks !",
      "votes": null
    },
    {
      "id": "937667",
      "postDate": "07/21/2020 05:48:15",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a> I didnt dig that deep to dct. This is a new knowledge to me. Thanks for the quick response.</p>\n<p>I tried shuffle without knowing this and my results were not good.</p>",
      "rawMarkdown": "Thanks @songwonho I didnt dig that deep to dct. This is a new knowledge to me. Thanks for the quick response.\n\nI tried shuffle without knowing this and my results were not good.",
      "votes": null
    },
    {
      "id": "937926",
      "postDate": "07/21/2020 08:39:46",
      "content": "<p><a href=\"/songwonho\">@songwonho</a> Congrats to you and your team..and Thanks for sharing.</p>",
      "rawMarkdown": "songwonho Congrats to you and your team..and Thanks for sharing.",
      "votes": null
    },
    {
      "id": "937960",
      "postDate": "07/21/2020 09:13:18",
      "content": "<p>Thanks for the detailed answer , I will the points in mind</p>",
      "rawMarkdown": "Thanks for the detailed answer , I will the points in mind",
      "votes": null
    },
    {
      "id": "938083",
      "postDate": "07/21/2020 10:35:40",
      "content": "<p>Thanks.</p>",
      "rawMarkdown": "Thanks.",
      "votes": null
    },
    {
      "id": "938103",
      "postDate": "07/21/2020 10:51:16",
      "content": "<p>Congratulations! Using GridShuffle is very clever. I thought about using CutMix but wasn't sure if it would work. </p>",
      "rawMarkdown": "Congratulations! Using GridShuffle is very clever. I thought about using CutMix but wasn't sure if it would work.",
      "votes": null
    },
    {
      "id": "938155",
      "postDate": "07/21/2020 11:26:57",
      "content": "<p>Thanks. Yes i agree with you. When we tried Cutmix, it was not work to boost cv scores. Congrats too!</p>",
      "rawMarkdown": "Thanks. Yes i agree with you. When we tried Cutmix, it was not work to boost cv scores. Congrats too!",
      "votes": null
    },
    {
      "id": "939015",
      "postDate": "07/22/2020 01:07:24",
      "content": "<p>That's some next level thinking on the augmentations and data. Good work!</p>",
      "rawMarkdown": "That's some next level thinking on the augmentations and data. Good work!",
      "votes": null
    },
    {
      "id": "941033",
      "postDate": "07/23/2020 06:04:21",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 937411,
      "author_name": "shwetank3",
      "author_url": "",
      "post_date": "07/21/2020 02:12:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a> ,</p>\n<p>Wonderful work. What I found most difficult was training took a lot of time especially B7 with pytorch. Can you help me with insight over your approach of tackling it. Was keras better than pytorch ? Did you use colab (GPU/TPU) to train ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 937469,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/21/2020 03:37:50",
          "content": "<p>thanks, We use pytorch. I think that large batch and large model would be great for this competition, so i tried to use TPU. However the competition was almost over, so I couldn't use it eventually with Reconstructed RGB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937560,
          "author_name": "pranshu29",
          "author_url": "",
          "post_date": "07/21/2020 04:43:00",
          "content": "<p>congrats <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a>  for the rank and thanks to share your pipeline but i want to ask something that why you have froze the backbone and fc layers ??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937574,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/21/2020 04:54:26",
          "content": "<p>Thanks. It`s a genius from giba. When we freeze backbone, we can train a model with a large batch size. Also, stego information is very small, we finetune the fully connected layer more through small learning rate.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937618,
      "author_name": "vishnurapps",
      "author_url": "",
      "post_date": "07/21/2020 05:29:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a> and team.</p>\n<p>I have a doubt regarding how grid shuffle and dropout help. When I checked using dct most of the information were in one part of the image and doing the shuffle disrupts the pattern. Could you please help me understand how it helped ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 937636,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/21/2020 05:39:16",
          "content": "<p>8x8 tiles was import in dct field. stego information was spread in 8x8 tiles and has a continuous distribution. Thus, while preserving this information, I proceeded with grid shuffle and grid dropout. My intention was that model could see stego information instead of object shape in image using 8x8 tile grid related augmentations. Details will be added to text soon!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937667,
          "author_name": "vishnurapps",
          "author_url": "",
          "post_date": "07/21/2020 05:48:15",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a> I didnt dig that deep to dct. This is a new knowledge to me. Thanks for the quick response.</p>\n<p>I tried shuffle without knowing this and my results were not good.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937659,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "07/21/2020 05:45:42",
      "content": "<p>Congrats to you and team on 8th place and thanks for sharing your solution <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a>!</p>",
      "votes": null,
      "replies": [
        {
          "id": 937664,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/21/2020 05:47:15",
          "content": "<p>Thanks !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 938103,
      "author_name": "vaillant",
      "author_url": "",
      "post_date": "07/21/2020 10:51:16",
      "content": "<p>Congratulations! Using GridShuffle is very clever. I thought about using CutMix but wasn't sure if it would work. </p>",
      "votes": null,
      "replies": [
        {
          "id": 938155,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/21/2020 11:26:57",
          "content": "<p>Thanks. Yes i agree with you. When we tried Cutmix, it was not work to boost cv scores. Congrats too!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937513,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "07/21/2020 04:10:06",
      "content": "<p>Thanks for a brief of the solution so early, I am very eager to learn although cv competitions are really computationally heavy and I don't have any gpus, can you give any advice/suggestions etc regarding this to a newbie like me\nThanks </p>",
      "votes": null,
      "replies": [
        {
          "id": 937589,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/21/2020 05:05:26",
          "content": "<p>Thanks. Nowadays cv competitions require huge resources. I think the fastest way is to buy a new powerful gpu, but it`s not easy. Here are some of method that have helped me.</p>\n\n<ol>\n<li><p>Use TPU\nWe can take advantage of a good TPU machine through Kaggle or Colab.</p></li>\n<li><p>Fast Experiment Settings\nIn the early stage, I set up an experimental environment that can see results quickly. For example, i tried a lot of data augmentations using efficientnet-b0 and large learning rate for fast convergence. Record good experiment results at this stage, and later apply good experiments through resources such as GCP or colab. What is important here is that the results in the fast experimental environment should align the results in the real experiment.</p></li>\n<li><p>Team-up\nIf we have a good idea, we can do it with colleagues with good equipment.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937621,
          "author_name": "vishnurapps",
          "author_url": "",
          "post_date": "07/21/2020 05:29:52",
          "content": "<p>Thank you for the insights. Its really helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937960,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "07/21/2020 09:13:18",
          "content": "<p>Thanks for the detailed answer , I will the points in mind</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937639,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "07/21/2020 05:39:56",
      "content": "<p>Thanks for sharing!\n😊 </p>",
      "votes": null,
      "replies": [
        {
          "id": 937649,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/21/2020 05:43:06",
          "content": "<p>thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937926,
      "author_name": "geekysaint",
      "author_url": "",
      "post_date": "07/21/2020 08:39:46",
      "content": "<p><a href=\"/songwonho\">@songwonho</a> Congrats to you and your team..and Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 938083,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/21/2020 10:35:40",
          "content": "<p>Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 939015,
      "author_name": "hooong",
      "author_url": "",
      "post_date": "07/22/2020 01:07:24",
      "content": "<p>That's some next level thinking on the augmentations and data. Good work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 941033,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/23/2020 06:04:21",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "937386": "# Before posting\nBefore posting, Thanks for hosting this great competition. And also, Thanks to our great teammates: @titericz, @youhanlee, @wonhyukahn, @bhjang.\n\nSorry for this brief solution, but details will be updated soon.\n\n# Pipeline\nOur pipelines are below:\n&gt; EDA -&gt; CV strategy -&gt; Data -&gt; Augmentations -&gt; Models -&gt; Ensemble\n\n# 1. EDA\n&gt; will be updated soon\n\n# 2. CV strategy\n&gt; - Stratified KFold based on quality factor and stego information insertion amount\n\nInformation insertion amount could be calculated below:\n&gt; a = difference between cover and stego to DCT coefficients per channel\nb = number of non-zero DCT coefficients per channel\ninformation insertion amount = a / b\n\nThen, After binning the information into 10, it was divided into k folds for each quality factor. At early stage we used 5 folds, but later 10 folds.\n\nBecause it took so long to train model, we didn't train all folds. One of the k folds was selected and used as hold-out.\n\n# 3. Data\n&gt; - Image: From RGB to ReconstructedRGB\n- Label: From binary to multi-label\n\nFor image, we use reconstructed RGB instead of RGB. When jpeg image is loaded using cv2, the small value is rounded and truncated. Since Stego information is a very fine information, it was important to preserve these values. Thus, We create an image without rounding and truncation using jpegio library.\n\nFor label, as mentioned earlier in discussions, multi-label classification was used rather than binary label. we used 10 targets (Cover, 3 quality factors * 3 stego algorithms).\n\n# 4. Augmentations\n&gt; - Flip, Rotate90, Cutout, GridShuffle, GridDropout\n\nI tried a lot for augmentation experiments.\nFirst most importantly, we remove augmentations that changes the original pixel value. Thus, we use **Flip, Rotate90, Cutout, GridShuffle, and GridDropout**. \n\nThe best result was GridShuffle. This augmentation came from thinking about the answer to the next question: How can our model focus on fine stego information, not the shape of the object in image? \nTo answer these question, We tried shuffle pixels then shuffle 8x8 tiles. In JPEG image, 8x8 tiles are important because DCT information is included in 8x8 tiles. Stego information is inserted at the DCT level, and since the DCT information is continuous within 8x8 tiles, We decided to mix these 8x8 tiles. \nIf the unit size of tile is set to 8x8, a total of 64*64 tiles are generated. Mixing all these tiles breaks the **texture information** along with the shape information, which hinders the model from being trained. Thus, we chose a suitable sized tile that is not too small and trained the model.\nThe GridDropout is similar.\n\ncode is below:\n\n* GridShuffle(Albumentations)\n&gt; grid_shuffle = OneOf([\n        RandomGridShuffle((2, 2), p=1.0),\n        RandomGridShuffle((2, 4), p=1.0),\n        RandomGridShuffle((2, 8), p=1.0),\n        RandomGridShuffle((2, 16), p=1.0),\n        RandomGridShuffle((2, 32), p=1.0),\n        RandomGridShuffle((4, 4), p=1.0),\n        RandomGridShuffle((4, 8), p=1.0),\n        RandomGridShuffle((4, 16), p=1.0),\n        RandomGridShuffle((4, 32), p=1.0),\n        RandomGridShuffle((8, 8), p=1.0),\n        RandomGridShuffle((8, 16), p=1.0),\n        RandomGridShuffle((8, 32), p=1.0),\n        RandomGridShuffle((16, 16), p=1.0),\n        RandomGridShuffle((16, 32), p=1.0),\n        RandomGridShuffle((32, 32), p=1.0),\n    ], p=0.7)\n\n\n# 5. Models\n&gt; - Framework: Pytorch\n- Backbone: Efficientnet b0, b0(double resolution), b3, b4, b7, resnext\n- Scheduler: ReduceOnplateau\n- Loss: CrossEntropy Loss\n- Optimizer: AdamP, AdamW\n- Epochs: 100+\n- Use TTA 8x\n\n# 6. ETC\n&gt; - Freeze backbone and finetune fc layer\n- Use all data as training\n- Flipped Stego Images\n\n# 7. Scores\n* Best Single Model in public leaderboard - Public LB: 945, Private LB: 925\n* Best Single Model in private leaderboard - Public LB: 940, Private LB: 928\n* Ensemble - Public LB: 944, Private LB: 929",
    "937411": "Hi @songwonho ,\n\nWonderful work. What I found most difficult was training took a lot of time especially B7 with pytorch. Can you help me with insight over your approach of tackling it. Was keras better than pytorch ? Did you use colab (GPU/TPU) to train ?",
    "937469": "thanks, We use pytorch. I think that large batch and large model would be great for this competition, so i tried to use TPU. However the competition was almost over, so I couldn't use it eventually with Reconstructed RGB.",
    "937513": "Thanks for a brief of the solution so early, I am very eager to learn although cv competitions are really computationally heavy and I don't have any gpus, can you give any advice/suggestions etc regarding this to a newbie like me\nThanks",
    "937560": "congrats @songwonho  for the rank and thanks to share your pipeline but i want to ask something that why you have froze the backbone and fc layers ??",
    "937574": "Thanks. It`s a genius from giba. When we freeze backbone, we can train a model with a large batch size. Also, stego information is very small, we finetune the fully connected layer more through small learning rate.",
    "937589": "Thanks. Nowadays cv competitions require huge resources. I think the fastest way is to buy a new powerful gpu, but it`s not easy. Here are some of method that have helped me.\n\n1. Use TPU\nWe can take advantage of a good TPU machine through Kaggle or Colab.\n\n2. Fast Experiment Settings\nIn the early stage, I set up an experimental environment that can see results quickly. For example, i tried a lot of data augmentations using efficientnet-b0 and large learning rate for fast convergence. Record good experiment results at this stage, and later apply good experiments through resources such as GCP or colab. What is important here is that the results in the fast experimental environment should align the results in the real experiment.\n\n3. Team-up\nIf we have a good idea, we can do it with colleagues with good equipment.",
    "937618": "Congrats @songwonho and team.\n\nI have a doubt regarding how grid shuffle and dropout help. When I checked using dct most of the information were in one part of the image and doing the shuffle disrupts the pattern. Could you please help me understand how it helped ?",
    "937621": "Thank you for the insights. Its really helpful.",
    "937636": "8x8 tiles was import in dct field. stego information was spread in 8x8 tiles and has a continuous distribution. Thus, while preserving this information, I proceeded with grid shuffle and grid dropout. My intention was that model could see stego information instead of object shape in image using 8x8 tile grid related augmentations. Details will be added to text soon!",
    "937639": "Thanks for sharing!\n😊",
    "937649": "thanks.",
    "937659": "Congrats to you and team on 8th place and thanks for sharing your solution @songwonho!",
    "937664": "Thanks !",
    "937667": "Thanks @songwonho I didnt dig that deep to dct. This is a new knowledge to me. Thanks for the quick response.\n\nI tried shuffle without knowing this and my results were not good.",
    "937926": "songwonho Congrats to you and your team..and Thanks for sharing.",
    "937960": "Thanks for the detailed answer , I will the points in mind",
    "938083": "Thanks.",
    "938103": "Congratulations! Using GridShuffle is very clever. I thought about using CutMix but wasn't sure if it would work.",
    "938155": "Thanks. Yes i agree with you. When we tried Cutmix, it was not work to boost cv scores. Congrats too!",
    "939015": "That's some next level thinking on the augmentations and data. Good work!",
    "941033": "Thanks!"
  },
  "source": "meta"
}