{
  "id": 155392,
  "title": "A few tips for easy start",
  "url": "/competitions/alaska2-image-steganalysis/discussion/155392",
  "author_name": "Eugene Khvedchenya",
  "post_date": "2020-06-01T14:52:44.288000",
  "votes": 157,
  "comment_count": 48,
  "views": 0,
  "content": "<p>The following is the set of tips that can make your start in this challenge a bit easier.</p>\n\n<ul>\n<li><strong>Don't resize</strong>. Any perturbation to original pixels will obliterate all hidden message (which is already a very weak signal). So don't resize, rotate or re-save images during training</li>\n<li>Augmentation with <strong>flips and 90-degree rotations are OK</strong>. </li>\n<li><strong>Crops are useless</strong>. In a sense, training on full-sized 512x512 tends to provide higher AUC score on all model architectures I've tried so far</li>\n<li><strong>Training on YCbCr colorspace is doubtful</strong> IMO. At the end of the day, YCbCr &lt;-&gt; RGB conversion is a linear combination and any CNN model will learn it with ease</li>\n<li><strong>Normalization is very important</strong>. Don't forget to do proper (at least ImageNet-like) normalization for your input data</li>\n<li><strong>Bronze with Resnet34 is doable</strong>.  It's faster to train than heavier models, but shows immediately if there is something wrong with your pipeline.</li>\n<li><strong>The way you read JPEGs is irrelevant</strong>. As long as you read images the same way for training &amp; testing you're good.</li>\n<li><strong>TTA helps</strong>, but up to some point. After 0.92 it worsens the score.</li>\n<li>You can solve this problem as binary- and multi-class classification. Multi-class seems to be a bit better, but it's not written in stone.</li>\n<li><strong>Bigger batch is better</strong>. I'm strongly suggesting to leverage fp16 training or if you know how to do model surgery - give a try to InplaceABN.</li>\n</ul>\n\n<p>I didn't explore hand-crafted features and using raw DCT coefficients so far, so this post will be updated as soon as I have some new insights for you. </p>",
  "messages": [
    {
      "id": 870198,
      "postDate": "2020-06-01T14:52:44.287Z",
      "content": "<p>The following is the set of tips that can make your start in this challenge a bit easier.</p>\n\n<ul>\n<li><strong>Don't resize</strong>. Any perturbation to original pixels will obliterate all hidden message (which is already a very weak signal). So don't resize, rotate or re-save images during training</li>\n<li>Augmentation with <strong>flips and 90-degree rotations are OK</strong>. </li>\n<li><strong>Crops are useless</strong>. In a sense, training on full-sized 512x512 tends to provide higher AUC score on all model architectures I've tried so far</li>\n<li><strong>Training on YCbCr colorspace is doubtful</strong> IMO. At the end of the day, YCbCr &lt;-&gt; RGB conversion is a linear combination and any CNN model will learn it with ease</li>\n<li><strong>Normalization is very important</strong>. Don't forget to do proper (at least ImageNet-like) normalization for your input data</li>\n<li><strong>Bronze with Resnet34 is doable</strong>.  It's faster to train than heavier models, but shows immediately if there is something wrong with your pipeline.</li>\n<li><strong>The way you read JPEGs is irrelevant</strong>. As long as you read images the same way for training &amp; testing you're good.</li>\n<li><strong>TTA helps</strong>, but up to some point. After 0.92 it worsens the score.</li>\n<li>You can solve this problem as binary- and multi-class classification. Multi-class seems to be a bit better, but it's not written in stone.</li>\n<li><strong>Bigger batch is better</strong>. I'm strongly suggesting to leverage fp16 training or if you know how to do model surgery - give a try to InplaceABN.</li>\n</ul>\n\n<p>I didn't explore hand-crafted features and using raw DCT coefficients so far, so this post will be updated as soon as I have some new insights for you. </p>",
      "rawMarkdown": "The following is the set of tips that can make your start in this challenge a bit easier.\n\n- **Don't resize**. Any perturbation to original pixels will obliterate all hidden message (which is already a very weak signal). So don't resize, rotate or re-save images during training\n- Augmentation with **flips and 90-degree rotations are OK**. \n- **Crops are useless**. In a sense, training on full-sized 512x512 tends to provide higher AUC score on all model architectures I've tried so far\n- **Training on YCbCr colorspace is doubtful** IMO. At the end of the day, YCbCr &lt;-&gt; RGB conversion is a linear combination and any CNN model will learn it with ease\n- **Normalization is very important**. Don't forget to do proper (at least ImageNet-like) normalization for your input data\n- **Bronze with Resnet34 is doable**.  It's faster to train than heavier models, but shows immediately if there is something wrong with your pipeline.\n- **The way you read JPEGs is irrelevant**. As long as you read images the same way for training &amp; testing you're good.\n- **TTA helps**, but up to some point. After 0.92 it worsens the score.\n- You can solve this problem as binary- and multi-class classification. Multi-class seems to be a bit better, but it's not written in stone.\n- **Bigger batch is better**. I'm strongly suggesting to leverage fp16 training or if you know how to do model surgery - give a try to InplaceABN.\n\nI didn't explore hand-crafted features and using raw DCT coefficients so far, so this post will be updated as soon as I have some new insights for you. ",
      "votes": 157
    },
    {
      "id": 871768,
      "postDate": "2020-06-02T16:14:49.510Z",
      "content": "<p>A lot of this post helped a lot of people.  I dropped from 18th to 75th over night!  But I like the community collaboration. \nStill a long way to go before the end of this.  My experiments confirm the following:</p>\n\n<p>1) Flips are a positive.  I have also had good luck with +/- 15% rotations as well.  A few fractions on CV corresponding with the same in LB</p>\n\n<p>2) DCT - I find no way to outperform JPEGs.  I abandoned this idea 2 weeks ago.</p>\n\n<p>3) Training on YCbCr, once you get the reads correct, is still under performing JPEGS, typically -0.01/0.02 LB</p>\n\n<p>4) No shock, but deeper networks do substantially better than shallower ones, at the top this will matter, but training time is much higher!</p>\n\n<p>5) TTA has always improved CV but been -0.0/0.01 on LB.  I don't use it anymore</p>\n\n<p>6) Multiclass is better in my experiments.  But once I got the improved performance, I stopped with the binary classifier experiments.</p>\n\n<p>7) Separate models by QFactor on JPEGs, has produced lower local CV.  I have not finished ensembling and will report back on these experiments when complete.</p>\n\n<p>8) Cosine Annealing takes longer but generates higher CV and LB.  Alternating epoch series between fixed LR and annealing LR seems to help as well, but I do not have the science to prove it.</p>\n\n<p>9) Val set hygiene is necessary. I am now matching CV and LB within +/-0.001 by balancing across all known factors.  My guess is the full test set looks a lot like the distribution of the train set.</p>",
      "rawMarkdown": "A lot of this post helped a lot of people.  I dropped from 18th to 75th over night!  But I like the community collaboration. \nStill a long way to go before the end of this.  My experiments confirm the following:\n\n1) Flips are a positive.  I have also had good luck with +/- 15% rotations as well.  A few fractions on CV corresponding with the same in LB\n\n2) DCT - I find no way to outperform JPEGs.  I abandoned this idea 2 weeks ago.\n\n3) Training on YCbCr, once you get the reads correct, is still under performing JPEGS, typically -0.01/0.02 LB\n\n4) No shock, but deeper networks do substantially better than shallower ones, at the top this will matter, but training time is much higher!\n\n5) TTA has always improved CV but been -0.0/0.01 on LB.  I don't use it anymore\n\n6) Multiclass is better in my experiments.  But once I got the improved performance, I stopped with the binary classifier experiments.\n\n7) Separate models by QFactor on JPEGs, has produced lower local CV.  I have not finished ensembling and will report back on these experiments when complete.\n\n8) Cosine Annealing takes longer but generates higher CV and LB.  Alternating epoch series between fixed LR and annealing LR seems to help as well, but I do not have the science to prove it.\n\n9) Val set hygiene is necessary. I am now matching CV and LB within +/-0.001 by balancing across all known factors.  My guess is the full test set looks a lot like the distribution of the train set.\n",
      "votes": 23,
      "replies": [
        {
          "id": 872707,
          "postDate": "2020-06-03T12:40:07.087Z",
          "content": "<p>Hi Brian and thanks for your hints!</p>\n\n<p>I'm quite surprised that rotations actually works in your case. It's an unexpected outcome (to me).</p>",
          "rawMarkdown": "Hi Brian and thanks for your hints!\n\nI'm quite surprised that rotations actually works in your case. It's an unexpected outcome (to me).\n\n",
          "votes": 1
        },
        {
          "id": 902427,
          "postDate": "2020-06-26T06:24:33.727Z",
          "content": "<p>Hi Brian, wondering how do you make a very good val set? I am facing the very large fluctuation of local val score vs Public LB</p>",
          "rawMarkdown": "Hi Brian, wondering how do you make a very good val set? I am facing the very large fluctuation of local val score vs Public LB",
          "votes": 1
        },
        {
          "id": 925730,
          "postDate": "2020-07-12T09:15:17.310Z",
          "content": "<p>Hi Brian, I really appreciate your advice. After trying both flipping and rotation, I gain improvement around 0.012 LB</p>",
          "rawMarkdown": "Hi Brian, I really appreciate your advice. After trying both flipping and rotation, I gain improvement around 0.012 LB",
          "votes": 1
        }
      ]
    },
    {
      "id": 896061,
      "postDate": "2020-06-21T19:41:47.930Z",
      "content": "<p>&gt; - Bigger batch is better. I'm strongly suggesting to leverage fp16 training or if you know how to do <strong>model surgery</strong> give a try to InplaceABN</p>\n\n<p>Here is the so called <strong>InPlaceABN model surgery</strong> for anyone interested:</p>\n\n<p><code>pip install git+https://github.com/mapillary/inplace_abn</code></p>\n\n<p>```\nfrom inplace_abn.abn import InPlaceABN</p>\n\n<p>def convert_layers(model, layer_type_old, layer_type_new, convert_weights=False):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            # recurse\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights)</p>\n\n<pre><code>    if type(module) == layer_type_old:\n        layer_old = module\n        layer_new = layer_type_new(module.num_features) \n\n        if convert_weights:\n            layer_new.weight = layer_old.weight\n            layer_new.bias = layer_old.bias\n\n        model._modules[name] = layer_new\n\nreturn model\n</code></pre>\n\n<p>convert_layers(net, nn.BatchNorm2d, InPlaceABN, True)</p>\n\n<p>net = net.cuda()\n```</p>\n\n<ul>\n<li>slightly modified convert_layers function from <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686\">this discussion topic</a></li>\n<li>created a <a href=\"https://www.kaggle.com/neongen/minimum-vram-footprint-gpu-baseline\">kernel</a> to showcase how to use <strong>InPlaceABN</strong> as well <strong>Apex</strong> to minimize gpu video memory footprint :). </li>\n</ul>",
      "rawMarkdown": "&gt; - Bigger batch is better. I'm strongly suggesting to leverage fp16 training or if you know how to do **model surgery** give a try to InplaceABN\n\nHere is the so called **InPlaceABN model surgery** for anyone interested:\n\n`pip install git+https://github.com/mapillary/inplace_abn`\n\n```\nfrom inplace_abn.abn import InPlaceABN\n\ndef convert_layers(model, layer_type_old, layer_type_new, convert_weights=False):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            # recurse\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights)\n\n        if type(module) == layer_type_old:\n            layer_old = module\n            layer_new = layer_type_new(module.num_features) \n\n            if convert_weights:\n                layer_new.weight = layer_old.weight\n                layer_new.bias = layer_old.bias\n\n            model._modules[name] = layer_new\n\n    return model\n\nconvert_layers(net, nn.BatchNorm2d, InPlaceABN, True)\n\nnet = net.cuda()\n```\n\n- slightly modified convert_layers function from [this discussion topic](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686)\n- created a [kernel](https://www.kaggle.com/neongen/minimum-vram-footprint-gpu-baseline) to showcase how to use **InPlaceABN** as well **Apex** to minimize gpu video memory footprint :). ",
      "votes": 15,
      "replies": [
        {
          "id": 899630,
          "postDate": "2020-06-24T10:44:02.483Z",
          "content": "<p>Don't you need to swap the pair BN + Act to ABN? Code above swaps BN to ABN but leaves activations as-is.</p>",
          "rawMarkdown": "Don't you need to swap the pair BN + Act to ABN? Code above swaps BN to ABN but leaves activations as-is.",
          "votes": 2
        },
        {
          "id": 899693,
          "postDate": "2020-06-24T11:49:29.100Z",
          "content": "<p>Good point. What we can do here is leaving activations as they are and passing activation = \"identity\" as argument for InPlaceABN. In this way we leave the original activations of the network to do their work in separate layers. The following modification to the code above is needed.</p>\n\n<p><code>layer_new = layer_type_new(module.num_features, activation=\"identity\")</code></p>",
          "rawMarkdown": "Good point. What we can do here is leaving activations as they are and passing activation = \"identity\" as argument for InPlaceABN. In this way we leave the original activations of the network to do their work in separate layers. The following modification to the code above is needed.\n\n`layer_new = layer_type_new(module.num_features, activation=\"identity\")`",
          "votes": 1
        },
        {
          "id": 899772,
          "postDate": "2020-06-24T12:39:02.090Z",
          "content": "<p>FWIW:\n```\nfrom inplace_abn.abn import InPlaceABN</p>\n\n<p>def convert_layers(model, layer_type_old, layer_type_new, convert_weights=False, p_name='',p_module=None):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights,p_name,p_module)</p>\n\n<pre><code>    if type(module) == layer_type_old and type(p_module) == Swish:\n        model._modules[p_name] = nn.Identity()\n        layer_old = module\n        layer_new = layer_type_new(module.num_features) if hasattr(module,'num_features') else layer_type_new()\n\n        if convert_weights:\n            layer_new.weight = layer_old.weight\n            layer_new.bias = layer_old.bias\n\n        model._modules[name] = layer_new\n    p_name, p_module = name, module\n\nreturn model\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "FWIW:\n```\nfrom inplace_abn.abn import InPlaceABN\n\ndef convert_layers(model, layer_type_old, layer_type_new, convert_weights=False, p_name='',p_module=None):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights,p_name,p_module)\n\n        if type(module) == layer_type_old and type(p_module) == Swish:\n            model._modules[p_name] = nn.Identity()\n            layer_old = module\n            layer_new = layer_type_new(module.num_features) if hasattr(module,'num_features') else layer_type_new()\n\n            if convert_weights:\n                layer_new.weight = layer_old.weight\n                layer_new.bias = layer_old.bias\n\n            model._modules[name] = layer_new\n        p_name, p_module = name, module\n\n    return model\n```",
          "votes": 2
        },
        {
          "id": 913173,
          "postDate": "2020-07-03T03:29:39.370Z",
          "content": "<p><img src=\"https://images.squarespace-cdn.com/content/548a09b7e4b09cb7481d6e1d/1438890878858-ZP1GD5FU2OMH7WQVZ5CY/?content-type=image%2Fjpeg\" alt=\"mvp\"></p>",
          "rawMarkdown": "![mvp](https://images.squarespace-cdn.com/content/548a09b7e4b09cb7481d6e1d/1438890878858-ZP1GD5FU2OMH7WQVZ5CY/?content-type=image%2Fjpeg)",
          "votes": 1
        },
        {
          "id": 918356,
          "postDate": "2020-07-07T07:49:44.420Z",
          "content": "<p>Are you guys getting comparable results with Inplace-BN? I am trying it for the first time and results are much worse. I wonder if it can be leaky-relu vs swish.</p>",
          "rawMarkdown": "Are you guys getting comparable results with Inplace-BN? I am trying it for the first time and results are much worse. I wonder if it can be leaky-relu vs swish."
        },
        {
          "id": 918405,
          "postDate": "2020-07-07T08:31:48.567Z",
          "content": "<p>I did not test it extensively as I m mostly exploiting apex training gains. I agree that worse results most probably come from  leaky-relu vs swish for Efficientnets, as you said. InPlaceABN offers the following activation functions: <code>relu, leaky_relu (default), elu, identity</code> \nSo there is no swish option available. Have you tried \"disabling\" activation inside Inplace-BN using <code>activation=\"identity\"</code> and do the activation in separate layers? </p>",
          "rawMarkdown": "I did not test it extensively as I m mostly exploiting apex training gains. I agree that worse results most probably come from  leaky-relu vs swish for Efficientnets, as you said. InPlaceABN offers the following activation functions: `relu, leaky_relu (default), elu, identity` \nSo there is no swish option available. Have you tried \"disabling\" activation inside Inplace-BN using `activation=\"identity\"` and do the activation in separate layers? "
        },
        {
          "id": 918441,
          "postDate": "2020-07-07T08:56:44.803Z",
          "content": "<p>Yes, this is more consistent, but runtime is slower actually because inplace bn is slower on itself than pytorch BN. Most speed gains come from switching from Swish to Leaky-Relu (less memory does not compensate).</p>",
          "rawMarkdown": "Yes, this is more consistent, but runtime is slower actually because inplace bn is slower on itself than pytorch BN. Most speed gains come from switching from Swish to Leaky-Relu (less memory does not compensate).",
          "votes": 1
        },
        {
          "id": 918485,
          "postDate": "2020-07-07T09:22:18Z",
          "content": "<p>Same here. I stopped using it.</p>",
          "rawMarkdown": "Same here. I stopped using it."
        },
        {
          "id": 918487,
          "postDate": "2020-07-07T09:22:50.380Z",
          "content": "<p>Using identity activation in inplace abn kills the whole idea of Inplace ABN :) of course no one is preventing from doing this, but it’s is useless, as you get no memory efficiency. It is also worth to mention that batchnorm in InplaceABN works slightly different to nn.BatchNorm2d. </p>",
          "rawMarkdown": "Using identity activation in inplace abn kills the whole idea of Inplace ABN :) of course no one is preventing from doing this, but it’s is useless, as you get no memory efficiency. It is also worth to mention that batchnorm in InplaceABN works slightly different to nn.BatchNorm2d. ",
          "votes": 5
        },
        {
          "id": 933032,
          "postDate": "2020-07-17T12:43:27.213Z",
          "content": "<p><a href=\"/bloodaxe\">@bloodaxe</a> <a href=\"/neongen\">@neongen</a> \nAbout Inplace ABN, \nWe've tried <a href=\"https://arxiv.org/abs/2003.13630\">TResNet</a> model which contains the <strong>In-Place Activated BatchNorm</strong> refinements to plain <strong>ResNet50</strong> design. The runtime was slow and convergence was much weaker. But we didn't more into it.</p>",
          "rawMarkdown": "@bloodaxe @neongen \nAbout Inplace ABN, \nWe've tried [TResNet](https://arxiv.org/abs/2003.13630) model which contains the **In-Place Activated BatchNorm** refinements to plain **ResNet50** design. The runtime was slow and convergence was much weaker. But we didn't more into it.\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 873517,
      "postDate": "2020-06-04T08:03:08.130Z",
      "content": "<p>For whatever reason, the image numbers aren't random. For example, in the first 10k images, the first 3k images are in Shanghai (two sets of photos actually), then 1.5k in Lille and Paris, then 1.5k of Brighton, UK, then 1.5k in France, then back to Brighton for 1k, then back to France for 1k. And those probably all taken by one person (someone who likes horses). While the Test images are themselves sequentially numbered from 1, the gaps in the Cover sequence may betray the location distribution of the actual test images. Whether this info is actually useful or not, hmmm ...</p>",
      "rawMarkdown": "For whatever reason, the image numbers aren't random. For example, in the first 10k images, the first 3k images are in Shanghai (two sets of photos actually), then 1.5k in Lille and Paris, then 1.5k of Brighton, UK, then 1.5k in France, then back to Brighton for 1k, then back to France for 1k. And those probably all taken by one person (someone who likes horses). While the Test images are themselves sequentially numbered from 1, the gaps in the Cover sequence may betray the location distribution of the actual test images. Whether this info is actually useful or not, hmmm ...",
      "votes": 11,
      "replies": [
        {
          "id": 907860,
          "postDate": "2020-06-30T07:52:51.253Z",
          "content": "<p>How did you extract location from those jpeg? I think the exif infomation have been removed 🤔 </p>",
          "rawMarkdown": "How did you extract location from those jpeg? I think the exif infomation have been removed 🤔 ",
          "votes": 1
        },
        {
          "id": 907901,
          "postDate": "2020-06-30T08:28:03.507Z",
          "content": "<p>Just visual observation, flipping through for an hour.</p>",
          "rawMarkdown": "Just visual observation, flipping through for an hour.",
          "votes": 1
        }
      ]
    },
    {
      "id": 873089,
      "postDate": "2020-06-03T19:20:38.673Z",
      "content": "<p>The most interesting aspect of this competition so far is that off the shelf pretrained models completely outperform previously hand engineered models found in the steganalysis literature. Even more surprising, perhaps, is that the existing literature hasn't tried to apply the best modern pretrained models to the task at hand. I suppose it's not novel enough to publish, but it sure works better. Let's see what happens before the competition is done ...</p>",
      "rawMarkdown": "The most interesting aspect of this competition so far is that off the shelf pretrained models completely outperform previously hand engineered models found in the steganalysis literature. Even more surprising, perhaps, is that the existing literature hasn't tried to apply the best modern pretrained models to the task at hand. I suppose it's not novel enough to publish, but it sure works better. Let's see what happens before the competition is done ...",
      "votes": 8,
      "replies": [
        {
          "id": 874351,
          "postDate": "2020-06-04T21:30:22.720Z",
          "content": "<p>maybe top N (i.e. N=10) open sourced solutions should collaboratively write an article :) . You can't be part of it if you don't provide the fully reproducible code. The author list is ordered by LB. A few wildcards could be given for nice visualizations, discussions or innovative ideas :) </p>",
          "rawMarkdown": "maybe top N (i.e. N=10) open sourced solutions should collaboratively write an article :) . You can't be part of it if you don't provide the fully reproducible code. The author list is ordered by LB. A few wildcards could be given for nice visualizations, discussions or innovative ideas :) "
        },
        {
          "id": 877311,
          "postDate": "2020-06-07T13:43:27.437Z",
          "content": "<p>The fact that pretrained models seem to outperform techniques find in literature is surprising me a lot. In my notebook I tried to follow common approaches like hand-crafted high-pass filters, truncating pixel values, using compression ratio as feature etc. Up to now the highest score I could achieve with these techniques was 85% which is significantly lower compared to just using a standard pretrained model</p>",
          "rawMarkdown": "The fact that pretrained models seem to outperform techniques find in literature is surprising me a lot. In my notebook I tried to follow common approaches like hand-crafted high-pass filters, truncating pixel values, using compression ratio as feature etc. Up to now the highest score I could achieve with these techniques was 85% which is significantly lower compared to just using a standard pretrained model"
        }
      ]
    },
    {
      "id": 872459,
      "postDate": "2020-06-03T07:57:12.773Z",
      "content": "<p>Wonderful tips, truly a piece of marvel.\nA few question however \npoint 4 - RGB vs YCbCr\n --&gt; Is this a though or do you have some results to support this claim ?\npoint 7 - Reading JPEG as pixels or DCT coefficients (at least this is how I interpreted \"the way you read JPEG\")\n --&gt; I agree that DCT is also a linear transformation, there the same argument holds \"it can be learnt\" \n --&gt; However, again, is this a though or do you have some results to support this claim ?</p>\n\n<p>A point that is not addressed, did you blend all images together (binary classification) or did you split the training / testing set in terms of JPEG quality factor / emebedding scheme (multi-class / targeted approach)</p>\n\n<p>Again, this post is terrific</p>",
      "rawMarkdown": "Wonderful tips, truly a piece of marvel.\nA few question however \npoint 4 - RGB vs YCbCr\n --&gt; Is this a though or do you have some results to support this claim ?\npoint 7 - Reading JPEG as pixels or DCT coefficients (at least this is how I interpreted \"the way you read JPEG\")\n --&gt; I agree that DCT is also a linear transformation, there the same argument holds \"it can be learnt\" \n --&gt; However, again, is this a though or do you have some results to support this claim ?\n\nA point that is not addressed, did you blend all images together (binary classification) or did you split the training / testing set in terms of JPEG quality factor / emebedding scheme (multi-class / targeted approach)\n\nAgain, this post is terrific\n",
      "votes": 3,
      "replies": [
        {
          "id": 872699,
          "postDate": "2020-06-03T12:36:51.563Z",
          "content": "<p>Hi <a href=\"/remicogranne\">@remicogranne</a> </p>\n\n<p>Thanks for your kind words, I have some past experience with digital image manipulation detection in JPEG image, maybe that helps </p>\n\n<blockquote>\n  <p>Is this a though or do you have some results to support this claim ?</p>\n</blockquote>\n\n<p>I did some ablation study by training a same model using RGB and YCBCR images as input (with a proper normalization for each colorspace) and observe no gain when using YCBCR. In fact, the latter was much worse. I can speculate this is due to the fact all pre-trained models assumes additive colorspace, so for YCBCR was used, a CNN had to re-learn all low-level filters to adapt to new domain. \nHowever, it's worth to notice I didn't run YCBCR training for a very long period of time. After 7-8 hours of training I stopped it, since it was clear at that point, that RGB input on the same pipeline gives a way better AUC. So I ditched exploration of YCBCR as for now.</p>\n\n<blockquote>\n  <p>--&gt; I agree that DCT is also a linear transformation, there the same argument holds \"it can be learnt\"\n  --&gt; However, again, is this a though or do you have some results to support this claim ?</p>\n</blockquote>\n\n<p>They are indeed can be learnt, in fact there are a few papers that exploit this approach.\nSo far, I did a few experiments with feeding DCT as input. Best results, I've got so far is 0.88 AUC, but I will continue exploration of this approach.</p>\n\n<blockquote>\n  <p>A point that is not addressed, did you blend all images together (binary classification) or did you split the training / testing set in terms of JPEG quality factor / emebedding scheme (multi-class / targeted approach)</p>\n</blockquote>\n\n<p>I partially mentioned it earlier. I actually train both - binary classifier and embedding-classifier. \nBoth demonstrate interesting insights, for instance how good a model can discriminate exact embedding scheme (left - train, right - validation)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1504864%2F35edb08dba3875cf64828cc5d06e5a66%2Fcm.png?generation=1591187783898834&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Hi @remicogranne \n\nThanks for your kind words, I have some past experience with digital image manipulation detection in JPEG image, maybe that helps \n\n&gt;  Is this a though or do you have some results to support this claim ?\n\nI did some ablation study by training a same model using RGB and YCBCR images as input (with a proper normalization for each colorspace) and observe no gain when using YCBCR. In fact, the latter was much worse. I can speculate this is due to the fact all pre-trained models assumes additive colorspace, so for YCBCR was used, a CNN had to re-learn all low-level filters to adapt to new domain. \nHowever, it's worth to notice I didn't run YCBCR training for a very long period of time. After 7-8 hours of training I stopped it, since it was clear at that point, that RGB input on the same pipeline gives a way better AUC. So I ditched exploration of YCBCR as for now.\n\n&gt; --&gt; I agree that DCT is also a linear transformation, there the same argument holds \"it can be learnt\"\n&gt; --&gt; However, again, is this a though or do you have some results to support this claim ?\n\nThey are indeed can be learnt, in fact there are a few papers that exploit this approach.\nSo far, I did a few experiments with feeding DCT as input. Best results, I've got so far is 0.88 AUC, but I will continue exploration of this approach.\n\n&gt; A point that is not addressed, did you blend all images together (binary classification) or did you split the training / testing set in terms of JPEG quality factor / emebedding scheme (multi-class / targeted approach)\n\nI partially mentioned it earlier. I actually train both - binary classifier and embedding-classifier. \nBoth demonstrate interesting insights, for instance how good a model can discriminate exact embedding scheme (left - train, right - validation)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1504864%2F35edb08dba3875cf64828cc5d06e5a66%2Fcm.png?generation=1591187783898834&amp;alt=media)\n\n\n",
          "votes": 12
        }
      ]
    },
    {
      "id": 870250,
      "postDate": "2020-06-01T15:30:17.250Z",
      "content": "<p>You are resetting the LB.  But Upvoted. :)</p>",
      "rawMarkdown": "You are resetting the LB.  But Upvoted. :)",
      "votes": 3,
      "replies": [
        {
          "id": 870365,
          "postDate": "2020-06-01T16:41:51.100Z",
          "content": "<p><a href=\"https://www.kaggle.com/shonenkov/train-inference-gpu-baseline\">https://www.kaggle.com/shonenkov/train-inference-gpu-baseline</a> already did this :)</p>",
          "rawMarkdown": "https://www.kaggle.com/shonenkov/train-inference-gpu-baseline already did this :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 870845,
      "postDate": "2020-06-02T01:30:08.957Z",
      "content": "<p>How many epochs do you think is worth training for? Let's say Efficient-Net-B0?</p>",
      "rawMarkdown": "How many epochs do you think is worth training for? Let's say Efficient-Net-B0?",
      "votes": 1,
      "replies": [
        {
          "id": 871288,
          "postDate": "2020-06-02T08:45:22.080Z",
          "content": "<p>until your validation loss is continuously decreasing. It depends on batch size and learning rate. As well as it depends on chosen optimizer.</p>",
          "rawMarkdown": "until your validation loss is continuously decreasing. It depends on batch size and learning rate. As well as it depends on chosen optimizer.",
          "votes": 3
        }
      ]
    },
    {
      "id": 870742,
      "postDate": "2020-06-01T21:48:00.480Z",
      "content": "<p>Upvoted! I got into the same conclusions and this will definitely help new people start.</p>",
      "rawMarkdown": "Upvoted! I got into the same conclusions and this will definitely help new people start.",
      "votes": 1
    },
    {
      "id": 883510,
      "postDate": "2020-06-12T17:43:24.837Z",
      "content": "<p>Thanks for the information :) </p>\n\n<p>I wonder do you train on all the data or say 10% of the data to iterate faster. Generally speaking more data is better but how different are they? </p>",
      "rawMarkdown": "Thanks for the information :) \n\nI wonder do you train on all the data or say 10% of the data to iterate faster. Generally speaking more data is better but how different are they? "
    },
    {
      "id": 883022,
      "postDate": "2020-06-12T09:48:54.347Z",
      "content": "<p>I saw the A.Resize(height=512, width=512, p=1.0) in the best score notebook from Alex. But why it needed?</p>",
      "rawMarkdown": "I saw the A.Resize(height=512, width=512, p=1.0) in the best score notebook from Alex. But why it needed?",
      "replies": [
        {
          "id": 883265,
          "postDate": "2020-06-12T13:51:27.727Z",
          "content": "<p>Perhaps, some left-over from other pipeline. In fact since images in this competition has size of 512, this transformation has zero effect.</p>",
          "rawMarkdown": "Perhaps, some left-over from other pipeline. In fact since images in this competition has size of 512, this transformation has zero effect."
        },
        {
          "id": 909406,
          "postDate": "2020-06-30T15:13:43.253Z",
          "content": "<p>For gpu, it is redundant. For tpu, you will get messy without it. <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu#760373\">Some explanation is here</a>. Working with tf.reshape instead of tf.image.resize is ok.</p>",
          "rawMarkdown": "For gpu, it is redundant. For tpu, you will get messy without it. [Some explanation is here](https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu#760373). Working with tf.reshape instead of tf.image.resize is ok.",
          "votes": 1
        }
      ]
    },
    {
      "id": 880475,
      "postDate": "2020-06-10T10:32:25.353Z",
      "content": "<p><a href=\"/bloodaxe\">@bloodaxe</a> that's huge tips. Thank you so much. :)</p>",
      "rawMarkdown": "@bloodaxe that's huge tips. Thank you so much. :)"
    },
    {
      "id": 875695,
      "postDate": "2020-06-06T05:11:09.633Z",
      "content": "<p>good</p>",
      "rawMarkdown": "good"
    },
    {
      "id": 874335,
      "postDate": "2020-06-04T21:08:06.200Z",
      "content": "<p>Thank you for very interesting insights!</p>",
      "rawMarkdown": "Thank you for very interesting insights!"
    },
    {
      "id": 871478,
      "postDate": "2020-06-02T11:47:14.550Z",
      "content": "<p>I just saved a screenshot of your tips! Thank you, <a href=\"/bloodaxe\">@bloodaxe</a>! Now I know why my model was scoring 60% accuracy; I was resizing the images 😅</p>",
      "rawMarkdown": "I just saved a screenshot of your tips! Thank you, @bloodaxe! Now I know why my model was scoring 60% accuracy; I was resizing the images 😅"
    },
    {
      "id": 871324,
      "postDate": "2020-06-02T09:12:36.263Z",
      "content": "<p>Thanks for saying \"Training on YCbCr colorspace is doubtful\". For a while I had trouble writing a proof for this but your post encouraged me to dig a little deeper and write <a href=\"https://www.kaggle.com/haveri/rgb-vs-ycbcr-this-may-be-why-it-does-not-matter\">this notebook</a>. There is indeed a linear relationship between RGB and YCbCr.</p>\n\n<p>I also thought using RGB (or YCbCr if one chooses to) along with DCT both 512x512 together could help. And the quantization tables too. Possibly:\n- a trainable pretrained model for RGB. Flatten the output.\n- a trainable pretrained model for DCT. Flatten the output.\n- the 2 quantization table of 8x8 matrix. Flatten them.</p>\n\n<p>Line them all one after the other for the final fully connected layers.</p>",
      "rawMarkdown": "Thanks for saying \"Training on YCbCr colorspace is doubtful\". For a while I had trouble writing a proof for this but your post encouraged me to dig a little deeper and write [this notebook](https://www.kaggle.com/haveri/rgb-vs-ycbcr-this-may-be-why-it-does-not-matter). There is indeed a linear relationship between RGB and YCbCr.\n\nI also thought using RGB (or YCbCr if one chooses to) along with DCT both 512x512 together could help. And the quantization tables too. Possibly:\n- a trainable pretrained model for RGB. Flatten the output.\n- a trainable pretrained model for DCT. Flatten the output.\n- the 2 quantization table of 8x8 matrix. Flatten them.\n\nLine them all one after the other for the final fully connected layers.",
      "replies": [
        {
          "id": 871360,
          "postDate": "2020-06-02T09:42:06.787Z",
          "content": "<p>So far, I was not quite successful with DCT-input models. Best result so far is ~0.88 AUC.  But I strongly agree that somehow we may benefit from using DCT signal and quantization information.</p>",
          "rawMarkdown": "So far, I was not quite successful with DCT-input models. Best result so far is ~0.88 AUC.  But I strongly agree that somehow we may benefit from using DCT signal and quantization information.",
          "votes": 1
        }
      ]
    },
    {
      "id": 870277,
      "postDate": "2020-06-01T15:48:47.687Z",
      "content": "<p>Thanks, Eugene, nice points. Were you able to get a good LB with binary classification? In my case I couldn't</p>",
      "rawMarkdown": "Thanks, Eugene, nice points. Were you able to get a good LB with binary classification? In my case I couldn't\n ",
      "replies": [
        {
          "id": 870367,
          "postDate": "2020-06-01T16:42:46.290Z",
          "content": "<p>Binary classifier performs a little (-0.001 LB) worse than multiclass. </p>",
          "rawMarkdown": "Binary classifier performs a little (-0.001 LB) worse than multiclass. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 870274,
      "postDate": "2020-06-01T15:45:45.237Z",
      "content": "<p>Thank you for very interesting insights!</p>",
      "rawMarkdown": "Thank you for very interesting insights!"
    },
    {
      "id": 924259,
      "postDate": "2020-07-11T10:56:57.940Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 871892,
      "postDate": "2020-06-02T17:58:50.147Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 871916,
          "postDate": "2020-06-02T18:13:30.210Z",
          "content": "<p>You can use curriculum learning, but i am pretty sure, that there is not much you can do with only kaggle gpu/tpu quota. </p>",
          "rawMarkdown": "You can use curriculum learning, but i am pretty sure, that there is not much you can do with only kaggle gpu/tpu quota. "
        }
      ]
    },
    {
      "id": 2045959,
      "postDate": "2022-11-27T20:11:10.417Z",
      "content": "<p>Thank you for sharing ! </p>",
      "rawMarkdown": "Thank you for sharing ! "
    },
    {
      "id": 887194,
      "postDate": "2020-06-15T14:29:16.303Z",
      "content": "<p>Very good tips.Thank you!</p>",
      "rawMarkdown": "Very good tips.Thank you!"
    },
    {
      "id": 873921,
      "postDate": "2020-06-04T14:19:22.283Z",
      "content": "<p>Thank you for very insights!</p>",
      "rawMarkdown": "Thank you for very insights!"
    },
    {
      "id": 873083,
      "postDate": "2020-06-03T19:10:43.657Z",
      "content": "<p>Thanks so much for such kind advice.</p>",
      "rawMarkdown": "Thanks so much for such kind advice."
    }
  ],
  "comments": [
    {
      "id": 871768,
      "author_name": "Brian Farrar",
      "author_url": "",
      "post_date": "2020-06-02T16:14:49.510000",
      "content": "<p>A lot of this post helped a lot of people.  I dropped from 18th to 75th over night!  But I like the community collaboration. \nStill a long way to go before the end of this.  My experiments confirm the following:</p>\n\n<p>1) Flips are a positive.  I have also had good luck with +/- 15% rotations as well.  A few fractions on CV corresponding with the same in LB</p>\n\n<p>2) DCT - I find no way to outperform JPEGs.  I abandoned this idea 2 weeks ago.</p>\n\n<p>3) Training on YCbCr, once you get the reads correct, is still under performing JPEGS, typically -0.01/0.02 LB</p>\n\n<p>4) No shock, but deeper networks do substantially better than shallower ones, at the top this will matter, but training time is much higher!</p>\n\n<p>5) TTA has always improved CV but been -0.0/0.01 on LB.  I don't use it anymore</p>\n\n<p>6) Multiclass is better in my experiments.  But once I got the improved performance, I stopped with the binary classifier experiments.</p>\n\n<p>7) Separate models by QFactor on JPEGs, has produced lower local CV.  I have not finished ensembling and will report back on these experiments when complete.</p>\n\n<p>8) Cosine Annealing takes longer but generates higher CV and LB.  Alternating epoch series between fixed LR and annealing LR seems to help as well, but I do not have the science to prove it.</p>\n\n<p>9) Val set hygiene is necessary. I am now matching CV and LB within +/-0.001 by balancing across all known factors.  My guess is the full test set looks a lot like the distribution of the train set.</p>",
      "votes": 23,
      "replies": [
        {
          "id": 872707,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-06-03T12:40:07.087000",
          "content": "<p>Hi Brian and thanks for your hints!</p>\n\n<p>I'm quite surprised that rotations actually works in your case. It's an unexpected outcome (to me).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 902427,
          "author_name": "Strideradu",
          "author_url": "",
          "post_date": "2020-06-26T06:24:33.727000",
          "content": "<p>Hi Brian, wondering how do you make a very good val set? I am facing the very large fluctuation of local val score vs Public LB</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 925730,
          "author_name": "Luck",
          "author_url": "",
          "post_date": "2020-07-12T09:15:17.310000",
          "content": "<p>Hi Brian, I really appreciate your advice. After trying both flipping and rotation, I gain improvement around 0.012 LB</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 896061,
      "author_name": "neongen",
      "author_url": "",
      "post_date": "2020-06-21T19:41:47.930000",
      "content": "<p>&gt; - Bigger batch is better. I'm strongly suggesting to leverage fp16 training or if you know how to do <strong>model surgery</strong> give a try to InplaceABN</p>\n\n<p>Here is the so called <strong>InPlaceABN model surgery</strong> for anyone interested:</p>\n\n<p><code>pip install git+https://github.com/mapillary/inplace_abn</code></p>\n\n<p>```\nfrom inplace_abn.abn import InPlaceABN</p>\n\n<p>def convert_layers(model, layer_type_old, layer_type_new, convert_weights=False):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            # recurse\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights)</p>\n\n<pre><code>    if type(module) == layer_type_old:\n        layer_old = module\n        layer_new = layer_type_new(module.num_features) \n\n        if convert_weights:\n            layer_new.weight = layer_old.weight\n            layer_new.bias = layer_old.bias\n\n        model._modules[name] = layer_new\n\nreturn model\n</code></pre>\n\n<p>convert_layers(net, nn.BatchNorm2d, InPlaceABN, True)</p>\n\n<p>net = net.cuda()\n```</p>\n\n<ul>\n<li>slightly modified convert_layers function from <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686\">this discussion topic</a></li>\n<li>created a <a href=\"https://www.kaggle.com/neongen/minimum-vram-footprint-gpu-baseline\">kernel</a> to showcase how to use <strong>InPlaceABN</strong> as well <strong>Apex</strong> to minimize gpu video memory footprint :). </li>\n</ul>",
      "votes": 15,
      "replies": [
        {
          "id": 899630,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2020-06-24T10:44:02.483000",
          "content": "<p>Don't you need to swap the pair BN + Act to ABN? Code above swaps BN to ABN but leaves activations as-is.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 899693,
          "author_name": "neongen",
          "author_url": "",
          "post_date": "2020-06-24T11:49:29.100000",
          "content": "<p>Good point. What we can do here is leaving activations as they are and passing activation = \"identity\" as argument for InPlaceABN. In this way we leave the original activations of the network to do their work in separate layers. The following modification to the code above is needed.</p>\n\n<p><code>layer_new = layer_type_new(module.num_features, activation=\"identity\")</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 899772,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2020-06-24T12:39:02.090000",
          "content": "<p>FWIW:\n```\nfrom inplace_abn.abn import InPlaceABN</p>\n\n<p>def convert_layers(model, layer_type_old, layer_type_new, convert_weights=False, p_name='',p_module=None):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights,p_name,p_module)</p>\n\n<pre><code>    if type(module) == layer_type_old and type(p_module) == Swish:\n        model._modules[p_name] = nn.Identity()\n        layer_old = module\n        layer_new = layer_type_new(module.num_features) if hasattr(module,'num_features') else layer_type_new()\n\n        if convert_weights:\n            layer_new.weight = layer_old.weight\n            layer_new.bias = layer_old.bias\n\n        model._modules[name] = layer_new\n    p_name, p_module = name, module\n\nreturn model\n</code></pre>\n\n<p>```</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 913173,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2020-07-03T03:29:39.370000",
          "content": "<p><img src=\"https://images.squarespace-cdn.com/content/548a09b7e4b09cb7481d6e1d/1438890878858-ZP1GD5FU2OMH7WQVZ5CY/?content-type=image%2Fjpeg\" alt=\"mvp\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 918356,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-07T07:49:44.420000",
          "content": "<p>Are you guys getting comparable results with Inplace-BN? I am trying it for the first time and results are much worse. I wonder if it can be leaky-relu vs swish.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 918405,
          "author_name": "neongen",
          "author_url": "",
          "post_date": "2020-07-07T08:31:48.567000",
          "content": "<p>I did not test it extensively as I m mostly exploiting apex training gains. I agree that worse results most probably come from  leaky-relu vs swish for Efficientnets, as you said. InPlaceABN offers the following activation functions: <code>relu, leaky_relu (default), elu, identity</code> \nSo there is no swish option available. Have you tried \"disabling\" activation inside Inplace-BN using <code>activation=\"identity\"</code> and do the activation in separate layers? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 918441,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-07-07T08:56:44.803000",
          "content": "<p>Yes, this is more consistent, but runtime is slower actually because inplace bn is slower on itself than pytorch BN. Most speed gains come from switching from Swish to Leaky-Relu (less memory does not compensate).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 918485,
          "author_name": "Andrés Miguel Torrubia Sáez",
          "author_url": "",
          "post_date": "2020-07-07T09:22:18",
          "content": "<p>Same here. I stopped using it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 918487,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-07-07T09:22:50.380000",
          "content": "<p>Using identity activation in inplace abn kills the whole idea of Inplace ABN :) of course no one is preventing from doing this, but it’s is useless, as you get no memory efficiency. It is also worth to mention that batchnorm in InplaceABN works slightly different to nn.BatchNorm2d. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 933032,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-07-17T12:43:27.213000",
          "content": "<p><a href=\"/bloodaxe\">@bloodaxe</a> <a href=\"/neongen\">@neongen</a> \nAbout Inplace ABN, \nWe've tried <a href=\"https://arxiv.org/abs/2003.13630\">TResNet</a> model which contains the <strong>In-Place Activated BatchNorm</strong> refinements to plain <strong>ResNet50</strong> design. The runtime was slow and convergence was much weaker. But we didn't more into it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 873517,
      "author_name": "robga",
      "author_url": "",
      "post_date": "2020-06-04T08:03:08.130000",
      "content": "<p>For whatever reason, the image numbers aren't random. For example, in the first 10k images, the first 3k images are in Shanghai (two sets of photos actually), then 1.5k in Lille and Paris, then 1.5k of Brighton, UK, then 1.5k in France, then back to Brighton for 1k, then back to France for 1k. And those probably all taken by one person (someone who likes horses). While the Test images are themselves sequentially numbered from 1, the gaps in the Cover sequence may betray the location distribution of the actual test images. Whether this info is actually useful or not, hmmm ...</p>",
      "votes": 11,
      "replies": [
        {
          "id": 907860,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-06-30T07:52:51.253000",
          "content": "<p>How did you extract location from those jpeg? I think the exif infomation have been removed 🤔 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 907901,
          "author_name": "robga",
          "author_url": "",
          "post_date": "2020-06-30T08:28:03.507000",
          "content": "<p>Just visual observation, flipping through for an hour.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 873089,
      "author_name": "robga",
      "author_url": "",
      "post_date": "2020-06-03T19:20:38.673000",
      "content": "<p>The most interesting aspect of this competition so far is that off the shelf pretrained models completely outperform previously hand engineered models found in the steganalysis literature. Even more surprising, perhaps, is that the existing literature hasn't tried to apply the best modern pretrained models to the task at hand. I suppose it's not novel enough to publish, but it sure works better. Let's see what happens before the competition is done ...</p>",
      "votes": 8,
      "replies": [
        {
          "id": 874351,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-04T21:30:22.720000",
          "content": "<p>maybe top N (i.e. N=10) open sourced solutions should collaboratively write an article :) . You can't be part of it if you don't provide the fully reproducible code. The author list is ordered by LB. A few wildcards could be given for nice visualizations, discussions or innovative ideas :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 877311,
          "author_name": "ManuelKraus89",
          "author_url": "",
          "post_date": "2020-06-07T13:43:27.437000",
          "content": "<p>The fact that pretrained models seem to outperform techniques find in literature is surprising me a lot. In my notebook I tried to follow common approaches like hand-crafted high-pass filters, truncating pixel values, using compression ratio as feature etc. Up to now the highest score I could achieve with these techniques was 85% which is significantly lower compared to just using a standard pretrained model</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 872459,
      "author_name": "Rémi Cogranne",
      "author_url": "",
      "post_date": "2020-06-03T07:57:12.773000",
      "content": "<p>Wonderful tips, truly a piece of marvel.\nA few question however \npoint 4 - RGB vs YCbCr\n --&gt; Is this a though or do you have some results to support this claim ?\npoint 7 - Reading JPEG as pixels or DCT coefficients (at least this is how I interpreted \"the way you read JPEG\")\n --&gt; I agree that DCT is also a linear transformation, there the same argument holds \"it can be learnt\" \n --&gt; However, again, is this a though or do you have some results to support this claim ?</p>\n\n<p>A point that is not addressed, did you blend all images together (binary classification) or did you split the training / testing set in terms of JPEG quality factor / emebedding scheme (multi-class / targeted approach)</p>\n\n<p>Again, this post is terrific</p>",
      "votes": 3,
      "replies": [
        {
          "id": 872699,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-06-03T12:36:51.563000",
          "content": "<p>Hi <a href=\"/remicogranne\">@remicogranne</a> </p>\n\n<p>Thanks for your kind words, I have some past experience with digital image manipulation detection in JPEG image, maybe that helps </p>\n\n<blockquote>\n  <p>Is this a though or do you have some results to support this claim ?</p>\n</blockquote>\n\n<p>I did some ablation study by training a same model using RGB and YCBCR images as input (with a proper normalization for each colorspace) and observe no gain when using YCBCR. In fact, the latter was much worse. I can speculate this is due to the fact all pre-trained models assumes additive colorspace, so for YCBCR was used, a CNN had to re-learn all low-level filters to adapt to new domain. \nHowever, it's worth to notice I didn't run YCBCR training for a very long period of time. After 7-8 hours of training I stopped it, since it was clear at that point, that RGB input on the same pipeline gives a way better AUC. So I ditched exploration of YCBCR as for now.</p>\n\n<blockquote>\n  <p>--&gt; I agree that DCT is also a linear transformation, there the same argument holds \"it can be learnt\"\n  --&gt; However, again, is this a though or do you have some results to support this claim ?</p>\n</blockquote>\n\n<p>They are indeed can be learnt, in fact there are a few papers that exploit this approach.\nSo far, I did a few experiments with feeding DCT as input. Best results, I've got so far is 0.88 AUC, but I will continue exploration of this approach.</p>\n\n<blockquote>\n  <p>A point that is not addressed, did you blend all images together (binary classification) or did you split the training / testing set in terms of JPEG quality factor / emebedding scheme (multi-class / targeted approach)</p>\n</blockquote>\n\n<p>I partially mentioned it earlier. I actually train both - binary classifier and embedding-classifier. \nBoth demonstrate interesting insights, for instance how good a model can discriminate exact embedding scheme (left - train, right - validation)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1504864%2F35edb08dba3875cf64828cc5d06e5a66%2Fcm.png?generation=1591187783898834&amp;alt=media\" alt=\"\"></p>",
          "votes": 12,
          "replies": []
        }
      ]
    },
    {
      "id": 870250,
      "author_name": "Johnny Lee",
      "author_url": "",
      "post_date": "2020-06-01T15:30:17.250000",
      "content": "<p>You are resetting the LB.  But Upvoted. :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 870365,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-06-01T16:41:51.100000",
          "content": "<p><a href=\"https://www.kaggle.com/shonenkov/train-inference-gpu-baseline\">https://www.kaggle.com/shonenkov/train-inference-gpu-baseline</a> already did this :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 870845,
      "author_name": "Ilia Ilmer",
      "author_url": "",
      "post_date": "2020-06-02T01:30:08.957000",
      "content": "<p>How many epochs do you think is worth training for? Let's say Efficient-Net-B0?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 871288,
          "author_name": "[ods.ai]Uladzimir Tumanau",
          "author_url": "",
          "post_date": "2020-06-02T08:45:22.080000",
          "content": "<p>until your validation loss is continuously decreasing. It depends on batch size and learning rate. As well as it depends on chosen optimizer.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 870742,
      "author_name": "GuillemDelgado",
      "author_url": "",
      "post_date": "2020-06-01T21:48:00.480000",
      "content": "<p>Upvoted! I got into the same conclusions and this will definitely help new people start.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 883510,
      "author_name": "Etudiant",
      "author_url": "",
      "post_date": "2020-06-12T17:43:24.837000",
      "content": "<p>Thanks for the information :) </p>\n\n<p>I wonder do you train on all the data or say 10% of the data to iterate faster. Generally speaking more data is better but how different are they? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 883022,
      "author_name": "Kirill Balakhonov",
      "author_url": "",
      "post_date": "2020-06-12T09:48:54.347000",
      "content": "<p>I saw the A.Resize(height=512, width=512, p=1.0) in the best score notebook from Alex. But why it needed?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 883265,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-06-12T13:51:27.727000",
          "content": "<p>Perhaps, some left-over from other pipeline. In fact since images in this competition has size of 512, this transformation has zero effect.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 909406,
          "author_name": "hklee",
          "author_url": "",
          "post_date": "2020-06-30T15:13:43.253000",
          "content": "<p>For gpu, it is redundant. For tpu, you will get messy without it. <a href=\"https://www.kaggle.com/mgornergoogle/getting-started-with-100-flowers-on-tpu#760373\">Some explanation is here</a>. Working with tf.reshape instead of tf.image.resize is ok.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 880475,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-06-10T10:32:25.353000",
      "content": "<p><a href=\"/bloodaxe\">@bloodaxe</a> that's huge tips. Thank you so much. :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 875695,
      "author_name": "Kallol Barai",
      "author_url": "",
      "post_date": "2020-06-06T05:11:09.633000",
      "content": "<p>good</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 874335,
      "author_name": "Saksham Varshney",
      "author_url": "",
      "post_date": "2020-06-04T21:08:06.200000",
      "content": "<p>Thank you for very interesting insights!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 871478,
      "author_name": "Tarun Paparaju",
      "author_url": "",
      "post_date": "2020-06-02T11:47:14.550000",
      "content": "<p>I just saved a screenshot of your tips! Thank you, <a href=\"/bloodaxe\">@bloodaxe</a>! Now I know why my model was scoring 60% accuracy; I was resizing the images 😅</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 871324,
      "author_name": "haveri",
      "author_url": "",
      "post_date": "2020-06-02T09:12:36.263000",
      "content": "<p>Thanks for saying \"Training on YCbCr colorspace is doubtful\". For a while I had trouble writing a proof for this but your post encouraged me to dig a little deeper and write <a href=\"https://www.kaggle.com/haveri/rgb-vs-ycbcr-this-may-be-why-it-does-not-matter\">this notebook</a>. There is indeed a linear relationship between RGB and YCbCr.</p>\n\n<p>I also thought using RGB (or YCbCr if one chooses to) along with DCT both 512x512 together could help. And the quantization tables too. Possibly:\n- a trainable pretrained model for RGB. Flatten the output.\n- a trainable pretrained model for DCT. Flatten the output.\n- the 2 quantization table of 8x8 matrix. Flatten them.</p>\n\n<p>Line them all one after the other for the final fully connected layers.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 871360,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-06-02T09:42:06.787000",
          "content": "<p>So far, I was not quite successful with DCT-input models. Best result so far is ~0.88 AUC.  But I strongly agree that somehow we may benefit from using DCT signal and quantization information.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 870277,
      "author_name": "Zhanseri Ikram",
      "author_url": "",
      "post_date": "2020-06-01T15:48:47.687000",
      "content": "<p>Thanks, Eugene, nice points. Were you able to get a good LB with binary classification? In my case I couldn't</p>",
      "votes": 0,
      "replies": [
        {
          "id": 870367,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-06-01T16:42:46.290000",
          "content": "<p>Binary classifier performs a little (-0.001 LB) worse than multiclass. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 870274,
      "author_name": "atfujita",
      "author_url": "",
      "post_date": "2020-06-01T15:45:45.237000",
      "content": "<p>Thank you for very interesting insights!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 924259,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T10:56:57.940000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 871892,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-02T17:58:50.147000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 871916,
          "author_name": "[ods.ai]Uladzimir Tumanau",
          "author_url": "",
          "post_date": "2020-06-02T18:13:30.210000",
          "content": "<p>You can use curriculum learning, but i am pretty sure, that there is not much you can do with only kaggle gpu/tpu quota. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2045959,
      "author_name": "ACL",
      "author_url": "",
      "post_date": "2022-11-27T20:11:10.417000",
      "content": "<p>Thank you for sharing ! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 887194,
      "author_name": "Md Selim Reza",
      "author_url": "",
      "post_date": "2020-06-15T14:29:16.303000",
      "content": "<p>Very good tips.Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 873921,
      "author_name": "Amit ",
      "author_url": "",
      "post_date": "2020-06-04T14:19:22.283000",
      "content": "<p>Thank you for very insights!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 873083,
      "author_name": "Ryan Hammang",
      "author_url": "",
      "post_date": "2020-06-03T19:10:43.657000",
      "content": "<p>Thanks so much for such kind advice.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "870198": "The following is the set of tips that can make your start in this challenge a bit easier.\n\n- **Don't resize**. Any perturbation to original pixels will obliterate all hidden message (which is already a very weak signal). So don't resize, rotate or re-save images during training\n- Augmentation with **flips and 90-degree rotations are OK**. \n- **Crops are useless**. In a sense, training on full-sized 512x512 tends to provide higher AUC score on all model architectures I've tried so far\n- **Training on YCbCr colorspace is doubtful** IMO. At the end of the day, YCbCr &lt;-&gt; RGB conversion is a linear combination and any CNN model will learn it with ease\n- **Normalization is very important**. Don't forget to do proper (at least ImageNet-like) normalization for your input data\n- **Bronze with Resnet34 is doable**.  It's faster to train than heavier models, but shows immediately if there is something wrong with your pipeline.\n- **The way you read JPEGs is irrelevant**. As long as you read images the same way for training &amp; testing you're good.\n- **TTA helps**, but up to some point. After 0.92 it worsens the score.\n- You can solve this problem as binary- and multi-class classification. Multi-class seems to be a bit better, but it's not written in stone.\n- **Bigger batch is better**. I'm strongly suggesting to leverage fp16 training or if you know how to do model surgery - give a try to InplaceABN.\n\nI didn't explore hand-crafted features and using raw DCT coefficients so far, so this post will be updated as soon as I have some new insights for you. ",
    "871768": "A lot of this post helped a lot of people.  I dropped from 18th to 75th over night!  But I like the community collaboration. \nStill a long way to go before the end of this.  My experiments confirm the following:\n\n1) Flips are a positive.  I have also had good luck with +/- 15% rotations as well.  A few fractions on CV corresponding with the same in LB\n\n2) DCT - I find no way to outperform JPEGs.  I abandoned this idea 2 weeks ago.\n\n3) Training on YCbCr, once you get the reads correct, is still under performing JPEGS, typically -0.01/0.02 LB\n\n4) No shock, but deeper networks do substantially better than shallower ones, at the top this will matter, but training time is much higher!\n\n5) TTA has always improved CV but been -0.0/0.01 on LB.  I don't use it anymore\n\n6) Multiclass is better in my experiments.  But once I got the improved performance, I stopped with the binary classifier experiments.\n\n7) Separate models by QFactor on JPEGs, has produced lower local CV.  I have not finished ensembling and will report back on these experiments when complete.\n\n8) Cosine Annealing takes longer but generates higher CV and LB.  Alternating epoch series between fixed LR and annealing LR seems to help as well, but I do not have the science to prove it.\n\n9) Val set hygiene is necessary. I am now matching CV and LB within +/-0.001 by balancing across all known factors.  My guess is the full test set looks a lot like the distribution of the train set.\n",
    "896061": "&gt; - Bigger batch is better. I'm strongly suggesting to leverage fp16 training or if you know how to do **model surgery** give a try to InplaceABN\n\nHere is the so called **InPlaceABN model surgery** for anyone interested:\n\n`pip install git+https://github.com/mapillary/inplace_abn`\n\n```\nfrom inplace_abn.abn import InPlaceABN\n\ndef convert_layers(model, layer_type_old, layer_type_new, convert_weights=False):\n    for name, module in reversed(model._modules.items()):\n        if len(list(module.children())) &gt; 0:\n            # recurse\n            model._modules[name] = convert_layers(module, layer_type_old, layer_type_new, convert_weights)\n\n        if type(module) == layer_type_old:\n            layer_old = module\n            layer_new = layer_type_new(module.num_features) \n\n            if convert_weights:\n                layer_new.weight = layer_old.weight\n                layer_new.bias = layer_old.bias\n\n            model._modules[name] = layer_new\n\n    return model\n\nconvert_layers(net, nn.BatchNorm2d, InPlaceABN, True)\n\nnet = net.cuda()\n```\n\n- slightly modified convert_layers function from [this discussion topic](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/104686)\n- created a [kernel](https://www.kaggle.com/neongen/minimum-vram-footprint-gpu-baseline) to showcase how to use **InPlaceABN** as well **Apex** to minimize gpu video memory footprint :). ",
    "873517": "For whatever reason, the image numbers aren't random. For example, in the first 10k images, the first 3k images are in Shanghai (two sets of photos actually), then 1.5k in Lille and Paris, then 1.5k of Brighton, UK, then 1.5k in France, then back to Brighton for 1k, then back to France for 1k. And those probably all taken by one person (someone who likes horses). While the Test images are themselves sequentially numbered from 1, the gaps in the Cover sequence may betray the location distribution of the actual test images. Whether this info is actually useful or not, hmmm ...",
    "873089": "The most interesting aspect of this competition so far is that off the shelf pretrained models completely outperform previously hand engineered models found in the steganalysis literature. Even more surprising, perhaps, is that the existing literature hasn't tried to apply the best modern pretrained models to the task at hand. I suppose it's not novel enough to publish, but it sure works better. Let's see what happens before the competition is done ...",
    "872459": "Wonderful tips, truly a piece of marvel.\nA few question however \npoint 4 - RGB vs YCbCr\n --&gt; Is this a though or do you have some results to support this claim ?\npoint 7 - Reading JPEG as pixels or DCT coefficients (at least this is how I interpreted \"the way you read JPEG\")\n --&gt; I agree that DCT is also a linear transformation, there the same argument holds \"it can be learnt\" \n --&gt; However, again, is this a though or do you have some results to support this claim ?\n\nA point that is not addressed, did you blend all images together (binary classification) or did you split the training / testing set in terms of JPEG quality factor / emebedding scheme (multi-class / targeted approach)\n\nAgain, this post is terrific\n",
    "870250": "You are resetting the LB.  But Upvoted. :)",
    "870845": "How many epochs do you think is worth training for? Let's say Efficient-Net-B0?",
    "870742": "Upvoted! I got into the same conclusions and this will definitely help new people start.",
    "883510": "Thanks for the information :) \n\nI wonder do you train on all the data or say 10% of the data to iterate faster. Generally speaking more data is better but how different are they? ",
    "883022": "I saw the A.Resize(height=512, width=512, p=1.0) in the best score notebook from Alex. But why it needed?",
    "880475": "@bloodaxe that's huge tips. Thank you so much. :)",
    "875695": "good",
    "874335": "Thank you for very interesting insights!",
    "871478": "I just saved a screenshot of your tips! Thank you, @bloodaxe! Now I know why my model was scoring 60% accuracy; I was resizing the images 😅",
    "871324": "Thanks for saying \"Training on YCbCr colorspace is doubtful\". For a while I had trouble writing a proof for this but your post encouraged me to dig a little deeper and write [this notebook](https://www.kaggle.com/haveri/rgb-vs-ycbcr-this-may-be-why-it-does-not-matter). There is indeed a linear relationship between RGB and YCbCr.\n\nI also thought using RGB (or YCbCr if one chooses to) along with DCT both 512x512 together could help. And the quantization tables too. Possibly:\n- a trainable pretrained model for RGB. Flatten the output.\n- a trainable pretrained model for DCT. Flatten the output.\n- the 2 quantization table of 8x8 matrix. Flatten them.\n\nLine them all one after the other for the final fully connected layers.",
    "870277": "Thanks, Eugene, nice points. Were you able to get a good LB with binary classification? In my case I couldn't\n ",
    "870274": "Thank you for very interesting insights!",
    "924259": "",
    "871892": "",
    "2045959": "Thank you for sharing ! ",
    "887194": "Very good tips.Thank you!",
    "873921": "Thank you for very insights!",
    "873083": "Thanks so much for such kind advice."
  }
}