{
  "id": 168546,
  "title": "[ABBA McCandless] - 2nd place solution overview ",
  "url": "/competitions/alaska2-image-steganalysis/writeups/abba-mccandless-abba-mccandless-2nd-place-solution",
  "author_name": "",
  "post_date": "2020-07-21T08:11:40.997Z",
  "votes": 75,
  "comment_count": 17,
  "views": 0,
  "content": "<p>A summary of our solution - 0.945/0.932 (Public/Private)</p>\n\n<p><strong>In short</strong>: Trust your CV</p>\n\n<p>I used 4 folds stratified by image size and quality factor with 20k holdout (5K Cover + 15K corresponding positive images). My team has a different split w/o folds but also with a holdout. We used the ranked average of second-level stacking (CatBoost and XGBoost) to blend individual models into ensembles and merge their predictions.</p>\n\n<p><strong>Models</strong>\nI used only B6/B7, and my teammates had B4/B5/B6/B7/MixNet plus SRNet and some hand-crafted features from \"classical\" steganalysis.\nI trained my models on RGB input (using cv2.imread) and then fine-tuned on non-rounded RGB images (by manually decoding DCT-&gt;YCbCr -&gt; RGB omitting rounding step). Lastly, I replaced Swish activation to Mish, which improved every model's score in my ensemble by +0.00044 on average.</p>\n\n<p><strong>Augmentations</strong>\nLike probably everyone, I've used D4 augmentations with coarse dropout for the training and D4 augmentations during the inference.</p>\n\n<p><strong>Losses</strong>\nI trained two heads for my models to predict binary (cover/stego) and multi-class (4 classes) outcomes. Based on early experiments, this showed to work well and speed up the convergence a little. I've played with different loss functions, including weighted BCE/CE, focal loss, and directly optimizing RoC AUC. But plain BCE/CE proved to work the best.</p>\n\n<p><strong>Ensembling</strong>\nFor ensembling I had tested many ideas, but the final solution was to train XGBoost classifier using predictions of my B6/B7 models. Second level stacking was done using holdout predictions. The feature matrix was 20000x380 made from binary and multi-class outputs, and all intermediate logits from TTA). I ran 5-fold cross-validation using groupkfold split on image_id to find the best hyperparameters for XGboost, and that was my ensemble.</p>\n\n<p>My team made a similar approach to their model zoo but trained CatBoost stacker. We selected our ensembles for the final blend based on individual CV scores and Spearman correlation. The lower correlation factor - the higher LB score on the public we had.</p>\n\n<p>To come to this solution, it took almost two months. There have been dozens of unsuccessful training runs and experiments on this way to gold. To name a few:\n* Training models on DCT input (Didn't even pass 0.92)\n* Training ResNet and DenseNet family (Barely reached 0.935 LB)\n* Training model to always have Cover + Stego image pair in batch \n* Direct Roc AUC optimization\n* Metric learning with ArcFace\n* Training with extra data (iStego 100K)\n* B0-B7, Heavier models performed better\n* Training stacking model on embeddings\n* Training on TPU. In general, it worked, but Kaggle TPU machine is limited in its CPU/RAM, so the training speed was deficient. </p>",
  "messages": [
    {
      "id": "937514",
      "postDate": "07/21/2020 04:10:20",
      "content": "<p>A summary of our solution - 0.945/0.932 (Public/Private)</p>\n\n<p><strong>In short</strong>: Trust your CV</p>\n\n<p>I used 4 folds stratified by image size and quality factor with 20k holdout (5K Cover + 15K corresponding positive images). My team has a different split w/o folds but also with a holdout. We used the ranked average of second-level stacking (CatBoost and XGBoost) to blend individual models into ensembles and merge their predictions.</p>\n\n<p><strong>Models</strong>\nI used only B6/B7, and my teammates had B4/B5/B6/B7/MixNet plus SRNet and some hand-crafted features from \"classical\" steganalysis.\nI trained my models on RGB input (using cv2.imread) and then fine-tuned on non-rounded RGB images (by manually decoding DCT-&gt;YCbCr -&gt; RGB omitting rounding step). Lastly, I replaced Swish activation to Mish, which improved every model's score in my ensemble by +0.00044 on average.</p>\n\n<p><strong>Augmentations</strong>\nLike probably everyone, I've used D4 augmentations with coarse dropout for the training and D4 augmentations during the inference.</p>\n\n<p><strong>Losses</strong>\nI trained two heads for my models to predict binary (cover/stego) and multi-class (4 classes) outcomes. Based on early experiments, this showed to work well and speed up the convergence a little. I've played with different loss functions, including weighted BCE/CE, focal loss, and directly optimizing RoC AUC. But plain BCE/CE proved to work the best.</p>\n\n<p><strong>Ensembling</strong>\nFor ensembling I had tested many ideas, but the final solution was to train XGBoost classifier using predictions of my B6/B7 models. Second level stacking was done using holdout predictions. The feature matrix was 20000x380 made from binary and multi-class outputs, and all intermediate logits from TTA). I ran 5-fold cross-validation using groupkfold split on image_id to find the best hyperparameters for XGboost, and that was my ensemble.</p>\n\n<p>My team made a similar approach to their model zoo but trained CatBoost stacker. We selected our ensembles for the final blend based on individual CV scores and Spearman correlation. The lower correlation factor - the higher LB score on the public we had.</p>\n\n<p>To come to this solution, it took almost two months. There have been dozens of unsuccessful training runs and experiments on this way to gold. To name a few:\n* Training models on DCT input (Didn't even pass 0.92)\n* Training ResNet and DenseNet family (Barely reached 0.935 LB)\n* Training model to always have Cover + Stego image pair in batch \n* Direct Roc AUC optimization\n* Metric learning with ArcFace\n* Training with extra data (iStego 100K)\n* B0-B7, Heavier models performed better\n* Training stacking model on embeddings\n* Training on TPU. In general, it worked, but Kaggle TPU machine is limited in its CPU/RAM, so the training speed was deficient. </p>",
      "rawMarkdown": "A summary of our solution - 0.945/0.932 (Public/Private)\n\n**In short**: Trust your CV\n\nI used 4 folds stratified by image size and quality factor with 20k holdout (5K Cover + 15K corresponding positive images). My team has a different split w/o folds but also with a holdout. We used the ranked average of second-level stacking (CatBoost and XGBoost) to blend individual models into ensembles and merge their predictions.\n\n**Models**\nI used only B6/B7, and my teammates had B4/B5/B6/B7/MixNet plus SRNet and some hand-crafted features from \"classical\" steganalysis.\nI trained my models on RGB input (using cv2.imread) and then fine-tuned on non-rounded RGB images (by manually decoding DCT-&gt;YCbCr -&gt; RGB omitting rounding step). Lastly, I replaced Swish activation to Mish, which improved every model's score in my ensemble by +0.00044 on average.\n\n**Augmentations**\nLike probably everyone, I've used D4 augmentations with coarse dropout for the training and D4 augmentations during the inference.\n\n**Losses**\nI trained two heads for my models to predict binary (cover/stego) and multi-class (4 classes) outcomes. Based on early experiments, this showed to work well and speed up the convergence a little. I've played with different loss functions, including weighted BCE/CE, focal loss, and directly optimizing RoC AUC. But plain BCE/CE proved to work the best.\n\n**Ensembling**\nFor ensembling I had tested many ideas, but the final solution was to train XGBoost classifier using predictions of my B6/B7 models. Second level stacking was done using holdout predictions. The feature matrix was 20000x380 made from binary and multi-class outputs, and all intermediate logits from TTA). I ran 5-fold cross-validation using groupkfold split on image_id to find the best hyperparameters for XGboost, and that was my ensemble.\n\nMy team made a similar approach to their model zoo but trained CatBoost stacker. We selected our ensembles for the final blend based on individual CV scores and Spearman correlation. The lower correlation factor - the higher LB score on the public we had.\n\nTo come to this solution, it took almost two months. There have been dozens of unsuccessful training runs and experiments on this way to gold. To name a few:\n* Training models on DCT input (Didn't even pass 0.92)\n* Training ResNet and DenseNet family (Barely reached 0.935 LB)\n* Training model to always have Cover + Stego image pair in batch \n* Direct Roc AUC optimization\n* Metric learning with ArcFace\n* Training with extra data (iStego 100K)\n* B0-B7, Heavier models performed better\n* Training stacking model on embeddings\n* Training on TPU. In general, it worked, but Kaggle TPU machine is limited in its CPU/RAM, so the training speed was deficient.",
      "votes": null
    },
    {
      "id": "937531",
      "postDate": "07/21/2020 04:24:28",
      "content": "<p>Congratulations on getting the second place.  <br>\nThanks for sharing the ideas. <br>\nOne question: What machine did you use for training? Did you train it offline or on Kaggle? </p>",
      "rawMarkdown": "Congratulations on getting the second place.  \nThanks for sharing the ideas. \nOne question: What machine did you use for training? Did you train it offline or on Kaggle?",
      "votes": null
    },
    {
      "id": "937536",
      "postDate": "07/21/2020 04:28:20",
      "content": "<p>Thank you!\nAround a year ago I built a 4x1080Ti dev. box that I used in this competition. \nFrom all challenges I've participated in, this particular challenge was probably the most demanding one in terms of hardware requirement. </p>",
      "rawMarkdown": "Thank you!\nAround a year ago I built a 4x1080Ti dev. box that I used in this competition. \nFrom all challenges I've participated in, this particular challenge was probably the most demanding one in terms of hardware requirement.",
      "votes": null
    },
    {
      "id": "937566",
      "postDate": "07/21/2020 04:48:56",
      "content": "<p>Great work! Congratulations :) </p>\n\n<p>First, I wonder the score before/after stacking.\nSecond, you said \"directly optimizing Roc AUC\", is there any reference or more details?</p>\n\n<p>Thanks for sharing.</p>",
      "rawMarkdown": "Great work! Congratulations :) \n\nFirst, I wonder the score before/after stacking.\nSecond, you said \"directly optimizing Roc AUC\", is there any reference or more details?\n\nThanks for sharing.",
      "votes": null
    },
    {
      "id": "937579",
      "postDate": "07/21/2020 04:55:51",
      "content": "<p>So, each fold of mine had bAUC and wAUC very close in range [0.931 - 0.933] on validation set  depending on the fold (No TTA).\nOn holdout with TTA they scored around 0.935-0.936 each. And with simple averaging - 0.9417 on holdout (0.9.\nWith XGBoost I was able to get 0.9422 CV (0.932 LB) on the holdout set.</p>\n\n<p>For optimizing RoC AUC I followed implementation and re-made it to PyTorch: <a href=\"https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py\">https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py</a></p>\n\n<p><code>``\nclass RocAucLoss(nn.Module):\n    \"\"\" ROC AUC Score.\n    Approximates the Area Under Curve score, using approximation based on\n    the Wilcoxon-Mann-Whitney U statistic.\n    Yan, L., Dodier, R., Mozer, M. C., &amp; Wolniewicz, R. (2003).\n    Optimizing Classifier Performance via an Approximation to the Wilcoxon-Mann-Whitney Statistic.\n    Measures overall performance for a full range of threshold levels.\n    Arguments:\n        y_pred:</code>Tensor<code>. Predicted values.\n        y_true:</code>Tensor` . Targets (labels), a probability distribution.\n    \"\"\"</p>\n\n<pre><code># https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py\ndef forward(self, y_pred, y_true):\n    eps = 1e-4\n    y_pred = torch.sigmoid(y_pred).clamp(eps, 1 - eps)\n    pos = y_pred[y_true == 1]\n    neg = y_pred[y_true == 0]\n\n    pos = torch.unsqueeze(pos, 0)\n    neg = torch.unsqueeze(neg, 1)\n\n    # original paper suggests performance is robust to exact parameter choice\n    gamma = 0.7\n    p = 2\n\n    difference = torch.zeros_like(pos * neg) + pos - neg - gamma\n    mask = difference &gt; 0\n    masked = difference.masked_fill(mask, 0)\n    return torch.mean(torch.pow(-masked, p))\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "So, each fold of mine had bAUC and wAUC very close in range [0.931 - 0.933] on validation set  depending on the fold (No TTA).\nOn holdout with TTA they scored around 0.935-0.936 each. And with simple averaging - 0.9417 on holdout (0.9.\nWith XGBoost I was able to get 0.9422 CV (0.932 LB) on the holdout set.\n\nFor optimizing RoC AUC I followed implementation and re-made it to PyTorch: https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py\n\n\n```\nclass RocAucLoss(nn.Module):\n    \"\"\" ROC AUC Score.\n    Approximates the Area Under Curve score, using approximation based on\n    the Wilcoxon-Mann-Whitney U statistic.\n    Yan, L., Dodier, R., Mozer, M. C., &amp; Wolniewicz, R. (2003).\n    Optimizing Classifier Performance via an Approximation to the Wilcoxon-Mann-Whitney Statistic.\n    Measures overall performance for a full range of threshold levels.\n    Arguments:\n        y_pred: `Tensor`. Predicted values.\n        y_true: `Tensor` . Targets (labels), a probability distribution.\n    \"\"\"\n\n    # https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py\n    def forward(self, y_pred, y_true):\n        eps = 1e-4\n        y_pred = torch.sigmoid(y_pred).clamp(eps, 1 - eps)\n        pos = y_pred[y_true == 1]\n        neg = y_pred[y_true == 0]\n\n        pos = torch.unsqueeze(pos, 0)\n        neg = torch.unsqueeze(neg, 1)\n\n        # original paper suggests performance is robust to exact parameter choice\n        gamma = 0.7\n        p = 2\n\n        difference = torch.zeros_like(pos * neg) + pos - neg - gamma\n        mask = difference &gt; 0\n        masked = difference.masked_fill(mask, 0)\n        return torch.mean(torch.pow(-masked, p))\n\n```",
      "votes": null
    },
    {
      "id": "937635",
      "postDate": "07/21/2020 05:39:11",
      "content": "<p>Thanks for sharing!\nAnd Congratulations 2nd place!</p>\n\n<p>You gave a lot of useful tips even at the beginning of the competition.\nThanks again! 👍 </p>",
      "rawMarkdown": "Thanks for sharing!\nAnd Congratulations 2nd place!\n\nYou gave a lot of useful tips even at the beginning of the competition.\nThanks again! 👍",
      "votes": null
    },
    {
      "id": "937642",
      "postDate": "07/21/2020 05:41:23",
      "content": "<p>Oh, thanks. It`s powerful. I learned a lot!</p>",
      "rawMarkdown": "Oh, thanks. It`s powerful. I learned a lot!",
      "votes": null
    },
    {
      "id": "937690",
      "postDate": "07/21/2020 06:01:06",
      "content": "<p>Hello there,</p>\n<p>Thanks for the quite detailed description and kudos for your second place ! :)</p>\n<p>Quick question however, you wrote <br>\n<em>\"I trained my models on RGB input (using cv2.imread) and then fine-tuned <strong>on non-rounded RGB images</strong> (by manually decoding DCT-&gt;YCbCr -&gt; <strong>DCT omitting rounding step</strong>\"</em><br>\nDo you YCbCr --&gt; RGB ? or  fine-tuned on DCT (it would case it would be non-sense starting on RGB and fine-tuning on somethg completely different)</p>\n<p>By the way, why didn' you use YCbCr directly (the data are indeed hidden in those channels) ?<br>\nAnd how did you made this process ? I provided a code to get YCbCr data but kagglers complainted it is way too slow ? Using cv2 built-in functions ? </p>\n<p>Other questions, did you blended al QF together ?</p>\n<p>Great post … and great achievement ;)</p>",
      "rawMarkdown": "Hello there,\n\nThanks for the quite detailed description and kudos for your second place ! :)\n\nQuick question however, you wrote \n*\"I trained my models on RGB input (using cv2.imread) and then fine-tuned **on non-rounded RGB images** (by manually decoding DCT-&gt;YCbCr -&gt; **DCT omitting rounding step**\"*\nDo you YCbCr --&gt; RGB ? or  fine-tuned on DCT (it would case it would be non-sense starting on RGB and fine-tuning on somethg completely different)\n\n\nBy the way, why didn' you use YCbCr directly (the data are indeed hidden in those channels) ?\nAnd how did you made this process ? I provided a code to get YCbCr data but kagglers complainted it is way too slow ? Using cv2 built-in functions ? \n\nOther questions, did you blended al QF together ?\n\nGreat post ... and great achievement ;)",
      "votes": null
    },
    {
      "id": "937705",
      "postDate": "07/21/2020 06:10:49",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> congrats</p>\n<p>I started this competition by reading your discussions in the <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/155392\" target=\"_blank\">forum </a>on how to get started. Waiting for your kernel 💯</p>",
      "rawMarkdown": "Hi @bloodaxe congrats\n\nI started this competition by reading your discussions in the [forum ](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/155392)on how to get started. Waiting for your kernel 💯",
      "votes": null
    },
    {
      "id": "937713",
      "postDate": "07/21/2020 06:15:26",
      "content": "<p>Congratulations on your second place!\nThe tips you posted on the discussion were very informative and helpful.\nI have one question: what are the settings for optimizer and lr? \nIn my case I got CV 0.930 LB 0.927 with a multiclassification of B5 + adamW + ReduceLROnPlateau(init_lr=1e-3, factor=0.4, patience=2). I'm surprised you can get 0.935 LB with ResNet and DenseNet.</p>",
      "rawMarkdown": "Congratulations on your second place!\nThe tips you posted on the discussion were very informative and helpful.\nI have one question: what are the settings for optimizer and lr? \nIn my case I got CV 0.930 LB 0.927 with a multiclassification of B5 + adamW + ReduceLROnPlateau(init_lr=1e-3, factor=0.4, patience=2). I'm surprised you can get 0.935 LB with ResNet and DenseNet.",
      "votes": null
    },
    {
      "id": "937747",
      "postDate": "07/21/2020 06:32:41",
      "content": "<p>Well summarized <a href=\"/bloodaxe\">@bloodaxe</a> It was a pleasure working with you! </p>\n\n<h3>A note on how to train SRNet</h3>\n\n<p>A few competitors have experienced difficulties in training SRNet, which is expectable since it's training from scratch (and not from a well trained ImageNet model), some finesse is needed to make it converge. \nWe first trained on QF75 with the largest BS we could afford (BS=64) following the training schedule described in the SRNet paper, then we fine-tuned to the other QFs. This was done because (if we assume fixed payload) QF75 is usually easier to detect. In this case, even if the payload widely varied, QF75 and QF90 were still easier than QF95.</p>\n\n<h3>Don't use cv2.imread</h3>\n\n<p>Unless you are willing to give up on free performance gains ;) We used a custom JPEG decoder which doesn't round/clip to [0,255], the gains can go up to 1% in wAUC (especially against JUNI). The forum had some examples of such decoders, we used our own implementation which takes advantage of numpy's stride_tricks for speed-ups (and jpegio of I/O). Some nets were first trained with cv2.imread then fine-tuned with the custom JPEG decoder.\nSRNet was trained on YCbCr, but the ImageNet pretrained models were trained on RGB (both non rounded). We believe that since ImageNet models were trained on RGB data, this would be more suitable. But in the end, it's a pretty easy linear transformation, I wouldn't be surprised if YCbCr worked as well.</p>\n\n<h3>A comparison of some of our models</h3>\n\n<p>Besides SRNet, the rest of the nets were trained on all QFs without any architecture modification. Note that the scatter shifts up when replacing Swish with Mish activation (c.f. original post), we couldn't do this for all our models in time ...\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1012681%2Fcf6a6c7f0c49fd6c34d158048b501084%2Fall_stego.png?generation=1595312025650266&amp;alt=media\" alt=\"\"></p>\n\n<h3>Good old hand crafted features</h3>\n\n<p>We stuck in some good old steganalysis hand crafted features (DCTR and JRM + FLD ensemble, which can both be found <a href=\"http://dde.binghamton.edu/download/feature_extractors/\">on our website</a>), they didn't perform the best (wAUC=0.850) but were added for the sake of diversity.</p>\n\n<h3>The stacking</h3>\n\n<p>We stacked our models and trained Catboost on holdout set (all QFs together), the QF was added as a categorical feature. Hyper parameters were tuned using <a href=\"https://scikit-optimize.github.io/stable/modules/generated/skopt.BayesSearchCV.html\">SKOPT BayesSearchCV</a>, but I'm sure a large enough grid search can do the job. Eugene Had a similar approach using xgboost.</p>\n\n<p><strong>Thanks to <a href=\"/remicogranne\">@remicogranne</a> , Patrick, Quentin and the Kaggle team for this competition! And thanks to everyone who made this competition so much fun!</strong>\n<em>More on our coming IEEE WIFS submission.</em></p>",
      "rawMarkdown": "Well summarized @bloodaxe It was a pleasure working with you! \n### A note on how to train SRNet\nA few competitors have experienced difficulties in training SRNet, which is expectable since it's training from scratch (and not from a well trained ImageNet model), some finesse is needed to make it converge. \nWe first trained on QF75 with the largest BS we could afford (BS=64) following the training schedule described in the SRNet paper, then we fine-tuned to the other QFs. This was done because (if we assume fixed payload) QF75 is usually easier to detect. In this case, even if the payload widely varied, QF75 and QF90 were still easier than QF95.\n### Don't use cv2.imread\nUnless you are willing to give up on free performance gains ;) We used a custom JPEG decoder which doesn't round/clip to [0,255], the gains can go up to 1% in wAUC (especially against JUNI). The forum had some examples of such decoders, we used our own implementation which takes advantage of numpy's stride_tricks for speed-ups (and jpegio of I/O). Some nets were first trained with cv2.imread then fine-tuned with the custom JPEG decoder.\nSRNet was trained on YCbCr, but the ImageNet pretrained models were trained on RGB (both non rounded). We believe that since ImageNet models were trained on RGB data, this would be more suitable. But in the end, it's a pretty easy linear transformation, I wouldn't be surprised if YCbCr worked as well.\n### A comparison of some of our models\nBesides SRNet, the rest of the nets were trained on all QFs without any architecture modification. Note that the scatter shifts up when replacing Swish with Mish activation (c.f. original post), we couldn't do this for all our models in time ...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1012681%2Fcf6a6c7f0c49fd6c34d158048b501084%2Fall_stego.png?generation=1595312025650266&amp;alt=media)\n### Good old hand crafted features\nWe stuck in some good old steganalysis hand crafted features (DCTR and JRM + FLD ensemble, which can both be found [on our website](http://dde.binghamton.edu/download/feature_extractors/)), they didn't perform the best (wAUC=0.850) but were added for the sake of diversity.\n### The stacking\nWe stacked our models and trained Catboost on holdout set (all QFs together), the QF was added as a categorical feature. Hyper parameters were tuned using [SKOPT BayesSearchCV](https://scikit-optimize.github.io/stable/modules/generated/skopt.BayesSearchCV.html), but I'm sure a large enough grid search can do the job. Eugene Had a similar approach using xgboost.\n\n**Thanks to @remicogranne , Patrick, Quentin and the Kaggle team for this competition! And thanks to everyone who made this competition so much fun!**\n*More on our coming IEEE WIFS submission.*",
      "votes": null
    },
    {
      "id": "937857",
      "postDate": "07/21/2020 07:50:42",
      "content": "<p>Thanks for sharing and a good explanation! 👍 </p>",
      "rawMarkdown": "Thanks for sharing and a good explanation! 👍",
      "votes": null
    },
    {
      "id": "937877",
      "postDate": "07/21/2020 08:10:55",
      "content": "<p>Thank you! \nI experimented with many hyper-parameters, and for me following combination worked quite good:\n- <code>SGD (lr=1e-2, wd=1e-4/1e-5)</code>, with cosine annealing to 1e-5.\n- <code>RAdam</code> (If starting from ImageNet) and <code>AdamW</code> (fine-tuning) with <code>lr=1e-4</code> and <code>wd=1e-2</code>.\n- <code>Ranger</code> with <code>flat + cos</code> schedule (For first half of the epochs LR is fixed, then decay with cosine to 0.01 of initial LR).\nB6/B7 models had dropout 0.5 after global average pooling.</p>",
      "rawMarkdown": "Thank you! \nI experimented with many hyper-parameters, and for me following combination worked quite good:\n- `SGD (lr=1e-2, wd=1e-4/1e-5)`, with cosine annealing to 1e-5.\n- `RAdam` (If starting from ImageNet) and `AdamW` (fine-tuning) with `lr=1e-4` and `wd=1e-2`.\n- `Ranger` with `flat + cos` schedule (For first half of the epochs LR is fixed, then decay with cosine to 0.01 of initial LR).\nB6/B7 models had dropout 0.5 after global average pooling.",
      "votes": null
    },
    {
      "id": "937886",
      "postDate": "07/21/2020 08:19:30",
      "content": "<p>Pardon for the typo in the summary. I tried to write asap and made a mistake in the description. \nThe conversion was DCT -&gt; YCbCr -&gt; RGB. But we computed non-rounded float32 RGB values. They were quite close to values one may get from <code>cv2.imread</code>, but this non-rounded input had an additional signal, which improved performance of every model we fine-tuned on non-rounded image input. Hope it clarifies our data pipeline.</p>\n\n<p>As for YCbCr, I initially experimented with it, but the CV/LB score was quite low, so I put on ice this approach in favor of RGB (known to work well) and test as many architectures I can. Same was with DCT input. I have almost identical approach as in 1st place solution, but haven't figured out how to incorporate embedding from DCT models in the second-level stacking. </p>\n\n<p>I included QF only in the second-level stacking model as one-hot vector. CV score with or without this feature was the same (maybe the difference was somewhere in 5 or 6 digit) but I decided to keep it just in case.</p>",
      "rawMarkdown": "Pardon for the typo in the summary. I tried to write asap and made a mistake in the description. \nThe conversion was DCT -&gt; YCbCr -&gt; RGB. But we computed non-rounded float32 RGB values. They were quite close to values one may get from `cv2.imread`, but this non-rounded input had an additional signal, which improved performance of every model we fine-tuned on non-rounded image input. Hope it clarifies our data pipeline.\n\nAs for YCbCr, I initially experimented with it, but the CV/LB score was quite low, so I put on ice this approach in favor of RGB (known to work well) and test as many architectures I can. Same was with DCT input. I have almost identical approach as in 1st place solution, but haven't figured out how to incorporate embedding from DCT models in the second-level stacking. \n\nI included QF only in the second-level stacking model as one-hot vector. CV score with or without this feature was the same (maybe the difference was somewhere in 5 or 6 digit) but I decided to keep it just in case.",
      "votes": null
    },
    {
      "id": "937889",
      "postDate": "07/21/2020 08:20:27",
      "content": "<p>Thank you for the detailed explanation!\nDid you plan to explain these parts later in the paper? 😊 </p>\n\n<p>Anyway, Congratulation 2nd place!</p>",
      "rawMarkdown": "Thank you for the detailed explanation!\nDid you plan to explain these parts later in the paper? 😊 \n\nAnyway, Congratulation 2nd place!",
      "votes": null
    },
    {
      "id": "937972",
      "postDate": "07/21/2020 09:23:28",
      "content": "<p>Congrats <a href=\"/bloodaxe\">@bloodaxe</a> Could you please share code. It always excellent learning going through your elegantly written code</p>",
      "rawMarkdown": "Congrats @bloodaxe Could you please share code. It always excellent learning going through your elegantly written code",
      "votes": null
    },
    {
      "id": "938102",
      "postDate": "07/21/2020 10:50:13",
      "content": "<p>\"Do you YCbCr --&gt; RGB ? or fine-tuned on DCT (it would case it would be non-sense starting on RGB and fine-tuning on somethg completely different)\"</p>\n\n<p>pretrain on RGB and later retrain on YCbCr actually works. I start with RGB and was unable to get good results. then i switch to YCbCr. to save time, i reuse rgb models as initiation when training ycbcr.</p>\n\n<p>I guess it can work because the gray channel in ycbcr is close to rgb, so the network just need to learn two more channels cb and cr. but the information in these two channels are not that strong</p>",
      "rawMarkdown": "\"Do you YCbCr --&gt; RGB ? or fine-tuned on DCT (it would case it would be non-sense starting on RGB and fine-tuning on somethg completely different)\"\n\npretrain on RGB and later retrain on YCbCr actually works. I start with RGB and was unable to get good results. then i switch to YCbCr. to save time, i reuse rgb models as initiation when training ycbcr.\n\nI guess it can work because the gray channel in ycbcr is close to rgb, so the network just need to learn two more channels cb and cr. but the information in these two channels are not that strong",
      "votes": null
    },
    {
      "id": "938671",
      "postDate": "07/21/2020 17:02:15",
      "content": "<p>Thanks for sharing your experince in detail <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> it is very illuminating!</p>",
      "rawMarkdown": "Thanks for sharing your experince in detail @bloodaxe it is very illuminating!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 937531,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "07/21/2020 04:24:28",
      "content": "<p>Congratulations on getting the second place.  <br>\nThanks for sharing the ideas. <br>\nOne question: What machine did you use for training? Did you train it offline or on Kaggle? </p>",
      "votes": null,
      "replies": [
        {
          "id": 937536,
          "author_name": "bloodaxe",
          "author_url": "",
          "post_date": "07/21/2020 04:28:20",
          "content": "<p>Thank you!\nAround a year ago I built a 4x1080Ti dev. box that I used in this competition. \nFrom all challenges I've participated in, this particular challenge was probably the most demanding one in terms of hardware requirement. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937690,
      "author_name": "remicogranne",
      "author_url": "",
      "post_date": "07/21/2020 06:01:06",
      "content": "<p>Hello there,</p>\n<p>Thanks for the quite detailed description and kudos for your second place ! :)</p>\n<p>Quick question however, you wrote <br>\n<em>\"I trained my models on RGB input (using cv2.imread) and then fine-tuned <strong>on non-rounded RGB images</strong> (by manually decoding DCT-&gt;YCbCr -&gt; <strong>DCT omitting rounding step</strong>\"</em><br>\nDo you YCbCr --&gt; RGB ? or  fine-tuned on DCT (it would case it would be non-sense starting on RGB and fine-tuning on somethg completely different)</p>\n<p>By the way, why didn' you use YCbCr directly (the data are indeed hidden in those channels) ?<br>\nAnd how did you made this process ? I provided a code to get YCbCr data but kagglers complainted it is way too slow ? Using cv2 built-in functions ? </p>\n<p>Other questions, did you blended al QF together ?</p>\n<p>Great post … and great achievement ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 937886,
          "author_name": "bloodaxe",
          "author_url": "",
          "post_date": "07/21/2020 08:19:30",
          "content": "<p>Pardon for the typo in the summary. I tried to write asap and made a mistake in the description. \nThe conversion was DCT -&gt; YCbCr -&gt; RGB. But we computed non-rounded float32 RGB values. They were quite close to values one may get from <code>cv2.imread</code>, but this non-rounded input had an additional signal, which improved performance of every model we fine-tuned on non-rounded image input. Hope it clarifies our data pipeline.</p>\n\n<p>As for YCbCr, I initially experimented with it, but the CV/LB score was quite low, so I put on ice this approach in favor of RGB (known to work well) and test as many architectures I can. Same was with DCT input. I have almost identical approach as in 1st place solution, but haven't figured out how to incorporate embedding from DCT models in the second-level stacking. </p>\n\n<p>I included QF only in the second-level stacking model as one-hot vector. CV score with or without this feature was the same (maybe the difference was somewhere in 5 or 6 digit) but I decided to keep it just in case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 938102,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/21/2020 10:50:13",
          "content": "<p>\"Do you YCbCr --&gt; RGB ? or fine-tuned on DCT (it would case it would be non-sense starting on RGB and fine-tuning on somethg completely different)\"</p>\n\n<p>pretrain on RGB and later retrain on YCbCr actually works. I start with RGB and was unable to get good results. then i switch to YCbCr. to save time, i reuse rgb models as initiation when training ycbcr.</p>\n\n<p>I guess it can work because the gray channel in ycbcr is close to rgb, so the network just need to learn two more channels cb and cr. but the information in these two channels are not that strong</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937705,
      "author_name": "vishnurapps",
      "author_url": "",
      "post_date": "07/21/2020 06:10:49",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> congrats</p>\n<p>I started this competition by reading your discussions in the <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/155392\" target=\"_blank\">forum </a>on how to get started. Waiting for your kernel 💯</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 938671,
      "author_name": "hgultekin",
      "author_url": "",
      "post_date": "07/21/2020 17:02:15",
      "content": "<p>Thanks for sharing your experince in detail <a href=\"https://www.kaggle.com/bloodaxe\" target=\"_blank\">@bloodaxe</a> it is very illuminating!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937566,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "07/21/2020 04:48:56",
      "content": "<p>Great work! Congratulations :) </p>\n\n<p>First, I wonder the score before/after stacking.\nSecond, you said \"directly optimizing Roc AUC\", is there any reference or more details?</p>\n\n<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 937579,
          "author_name": "bloodaxe",
          "author_url": "",
          "post_date": "07/21/2020 04:55:51",
          "content": "<p>So, each fold of mine had bAUC and wAUC very close in range [0.931 - 0.933] on validation set  depending on the fold (No TTA).\nOn holdout with TTA they scored around 0.935-0.936 each. And with simple averaging - 0.9417 on holdout (0.9.\nWith XGBoost I was able to get 0.9422 CV (0.932 LB) on the holdout set.</p>\n\n<p>For optimizing RoC AUC I followed implementation and re-made it to PyTorch: <a href=\"https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py\">https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py</a></p>\n\n<p><code>``\nclass RocAucLoss(nn.Module):\n    \"\"\" ROC AUC Score.\n    Approximates the Area Under Curve score, using approximation based on\n    the Wilcoxon-Mann-Whitney U statistic.\n    Yan, L., Dodier, R., Mozer, M. C., &amp; Wolniewicz, R. (2003).\n    Optimizing Classifier Performance via an Approximation to the Wilcoxon-Mann-Whitney Statistic.\n    Measures overall performance for a full range of threshold levels.\n    Arguments:\n        y_pred:</code>Tensor<code>. Predicted values.\n        y_true:</code>Tensor` . Targets (labels), a probability distribution.\n    \"\"\"</p>\n\n<pre><code># https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py\ndef forward(self, y_pred, y_true):\n    eps = 1e-4\n    y_pred = torch.sigmoid(y_pred).clamp(eps, 1 - eps)\n    pos = y_pred[y_true == 1]\n    neg = y_pred[y_true == 0]\n\n    pos = torch.unsqueeze(pos, 0)\n    neg = torch.unsqueeze(neg, 1)\n\n    # original paper suggests performance is robust to exact parameter choice\n    gamma = 0.7\n    p = 2\n\n    difference = torch.zeros_like(pos * neg) + pos - neg - gamma\n    mask = difference &gt; 0\n    masked = difference.masked_fill(mask, 0)\n    return torch.mean(torch.pow(-masked, p))\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937642,
          "author_name": "songwonho",
          "author_url": "",
          "post_date": "07/21/2020 05:41:23",
          "content": "<p>Oh, thanks. It`s powerful. I learned a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937635,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "07/21/2020 05:39:11",
      "content": "<p>Thanks for sharing!\nAnd Congratulations 2nd place!</p>\n\n<p>You gave a lot of useful tips even at the beginning of the competition.\nThanks again! 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937713,
      "author_name": "ajtryt2",
      "author_url": "",
      "post_date": "07/21/2020 06:15:26",
      "content": "<p>Congratulations on your second place!\nThe tips you posted on the discussion were very informative and helpful.\nI have one question: what are the settings for optimizer and lr? \nIn my case I got CV 0.930 LB 0.927 with a multiclassification of B5 + adamW + ReduceLROnPlateau(init_lr=1e-3, factor=0.4, patience=2). I'm surprised you can get 0.935 LB with ResNet and DenseNet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 937877,
          "author_name": "bloodaxe",
          "author_url": "",
          "post_date": "07/21/2020 08:10:55",
          "content": "<p>Thank you! \nI experimented with many hyper-parameters, and for me following combination worked quite good:\n- <code>SGD (lr=1e-2, wd=1e-4/1e-5)</code>, with cosine annealing to 1e-5.\n- <code>RAdam</code> (If starting from ImageNet) and <code>AdamW</code> (fine-tuning) with <code>lr=1e-4</code> and <code>wd=1e-2</code>.\n- <code>Ranger</code> with <code>flat + cos</code> schedule (For first half of the epochs LR is fixed, then decay with cosine to 0.01 of initial LR).\nB6/B7 models had dropout 0.5 after global average pooling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 937889,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "07/21/2020 08:20:27",
          "content": "<p>Thank you for the detailed explanation!\nDid you plan to explain these parts later in the paper? 😊 </p>\n\n<p>Anyway, Congratulation 2nd place!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937747,
      "author_name": "yousfi",
      "author_url": "",
      "post_date": "07/21/2020 06:32:41",
      "content": "<p>Well summarized <a href=\"/bloodaxe\">@bloodaxe</a> It was a pleasure working with you! </p>\n\n<h3>A note on how to train SRNet</h3>\n\n<p>A few competitors have experienced difficulties in training SRNet, which is expectable since it's training from scratch (and not from a well trained ImageNet model), some finesse is needed to make it converge. \nWe first trained on QF75 with the largest BS we could afford (BS=64) following the training schedule described in the SRNet paper, then we fine-tuned to the other QFs. This was done because (if we assume fixed payload) QF75 is usually easier to detect. In this case, even if the payload widely varied, QF75 and QF90 were still easier than QF95.</p>\n\n<h3>Don't use cv2.imread</h3>\n\n<p>Unless you are willing to give up on free performance gains ;) We used a custom JPEG decoder which doesn't round/clip to [0,255], the gains can go up to 1% in wAUC (especially against JUNI). The forum had some examples of such decoders, we used our own implementation which takes advantage of numpy's stride_tricks for speed-ups (and jpegio of I/O). Some nets were first trained with cv2.imread then fine-tuned with the custom JPEG decoder.\nSRNet was trained on YCbCr, but the ImageNet pretrained models were trained on RGB (both non rounded). We believe that since ImageNet models were trained on RGB data, this would be more suitable. But in the end, it's a pretty easy linear transformation, I wouldn't be surprised if YCbCr worked as well.</p>\n\n<h3>A comparison of some of our models</h3>\n\n<p>Besides SRNet, the rest of the nets were trained on all QFs without any architecture modification. Note that the scatter shifts up when replacing Swish with Mish activation (c.f. original post), we couldn't do this for all our models in time ...\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1012681%2Fcf6a6c7f0c49fd6c34d158048b501084%2Fall_stego.png?generation=1595312025650266&amp;alt=media\" alt=\"\"></p>\n\n<h3>Good old hand crafted features</h3>\n\n<p>We stuck in some good old steganalysis hand crafted features (DCTR and JRM + FLD ensemble, which can both be found <a href=\"http://dde.binghamton.edu/download/feature_extractors/\">on our website</a>), they didn't perform the best (wAUC=0.850) but were added for the sake of diversity.</p>\n\n<h3>The stacking</h3>\n\n<p>We stacked our models and trained Catboost on holdout set (all QFs together), the QF was added as a categorical feature. Hyper parameters were tuned using <a href=\"https://scikit-optimize.github.io/stable/modules/generated/skopt.BayesSearchCV.html\">SKOPT BayesSearchCV</a>, but I'm sure a large enough grid search can do the job. Eugene Had a similar approach using xgboost.</p>\n\n<p><strong>Thanks to <a href=\"/remicogranne\">@remicogranne</a> , Patrick, Quentin and the Kaggle team for this competition! And thanks to everyone who made this competition so much fun!</strong>\n<em>More on our coming IEEE WIFS submission.</em></p>",
      "votes": null,
      "replies": [
        {
          "id": 937857,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "07/21/2020 07:50:42",
          "content": "<p>Thanks for sharing and a good explanation! 👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937972,
      "author_name": "kalyanpichuka",
      "author_url": "",
      "post_date": "07/21/2020 09:23:28",
      "content": "<p>Congrats <a href=\"/bloodaxe\">@bloodaxe</a> Could you please share code. It always excellent learning going through your elegantly written code</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "937514": "A summary of our solution - 0.945/0.932 (Public/Private)\n\n**In short**: Trust your CV\n\nI used 4 folds stratified by image size and quality factor with 20k holdout (5K Cover + 15K corresponding positive images). My team has a different split w/o folds but also with a holdout. We used the ranked average of second-level stacking (CatBoost and XGBoost) to blend individual models into ensembles and merge their predictions.\n\n**Models**\nI used only B6/B7, and my teammates had B4/B5/B6/B7/MixNet plus SRNet and some hand-crafted features from \"classical\" steganalysis.\nI trained my models on RGB input (using cv2.imread) and then fine-tuned on non-rounded RGB images (by manually decoding DCT-&gt;YCbCr -&gt; RGB omitting rounding step). Lastly, I replaced Swish activation to Mish, which improved every model's score in my ensemble by +0.00044 on average.\n\n**Augmentations**\nLike probably everyone, I've used D4 augmentations with coarse dropout for the training and D4 augmentations during the inference.\n\n**Losses**\nI trained two heads for my models to predict binary (cover/stego) and multi-class (4 classes) outcomes. Based on early experiments, this showed to work well and speed up the convergence a little. I've played with different loss functions, including weighted BCE/CE, focal loss, and directly optimizing RoC AUC. But plain BCE/CE proved to work the best.\n\n**Ensembling**\nFor ensembling I had tested many ideas, but the final solution was to train XGBoost classifier using predictions of my B6/B7 models. Second level stacking was done using holdout predictions. The feature matrix was 20000x380 made from binary and multi-class outputs, and all intermediate logits from TTA). I ran 5-fold cross-validation using groupkfold split on image_id to find the best hyperparameters for XGboost, and that was my ensemble.\n\nMy team made a similar approach to their model zoo but trained CatBoost stacker. We selected our ensembles for the final blend based on individual CV scores and Spearman correlation. The lower correlation factor - the higher LB score on the public we had.\n\nTo come to this solution, it took almost two months. There have been dozens of unsuccessful training runs and experiments on this way to gold. To name a few:\n* Training models on DCT input (Didn't even pass 0.92)\n* Training ResNet and DenseNet family (Barely reached 0.935 LB)\n* Training model to always have Cover + Stego image pair in batch \n* Direct Roc AUC optimization\n* Metric learning with ArcFace\n* Training with extra data (iStego 100K)\n* B0-B7, Heavier models performed better\n* Training stacking model on embeddings\n* Training on TPU. In general, it worked, but Kaggle TPU machine is limited in its CPU/RAM, so the training speed was deficient.",
    "937531": "Congratulations on getting the second place.  \nThanks for sharing the ideas. \nOne question: What machine did you use for training? Did you train it offline or on Kaggle?",
    "937536": "Thank you!\nAround a year ago I built a 4x1080Ti dev. box that I used in this competition. \nFrom all challenges I've participated in, this particular challenge was probably the most demanding one in terms of hardware requirement.",
    "937566": "Great work! Congratulations :) \n\nFirst, I wonder the score before/after stacking.\nSecond, you said \"directly optimizing Roc AUC\", is there any reference or more details?\n\nThanks for sharing.",
    "937579": "So, each fold of mine had bAUC and wAUC very close in range [0.931 - 0.933] on validation set  depending on the fold (No TTA).\nOn holdout with TTA they scored around 0.935-0.936 each. And with simple averaging - 0.9417 on holdout (0.9.\nWith XGBoost I was able to get 0.9422 CV (0.932 LB) on the holdout set.\n\nFor optimizing RoC AUC I followed implementation and re-made it to PyTorch: https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py\n\n\n```\nclass RocAucLoss(nn.Module):\n    \"\"\" ROC AUC Score.\n    Approximates the Area Under Curve score, using approximation based on\n    the Wilcoxon-Mann-Whitney U statistic.\n    Yan, L., Dodier, R., Mozer, M. C., &amp; Wolniewicz, R. (2003).\n    Optimizing Classifier Performance via an Approximation to the Wilcoxon-Mann-Whitney Statistic.\n    Measures overall performance for a full range of threshold levels.\n    Arguments:\n        y_pred: `Tensor`. Predicted values.\n        y_true: `Tensor` . Targets (labels), a probability distribution.\n    \"\"\"\n\n    # https://github.com/tflearn/tflearn/blob/5a674b7f7d70064c811cbd98c4a41a17893d44ee/tflearn/objectives.py\n    def forward(self, y_pred, y_true):\n        eps = 1e-4\n        y_pred = torch.sigmoid(y_pred).clamp(eps, 1 - eps)\n        pos = y_pred[y_true == 1]\n        neg = y_pred[y_true == 0]\n\n        pos = torch.unsqueeze(pos, 0)\n        neg = torch.unsqueeze(neg, 1)\n\n        # original paper suggests performance is robust to exact parameter choice\n        gamma = 0.7\n        p = 2\n\n        difference = torch.zeros_like(pos * neg) + pos - neg - gamma\n        mask = difference &gt; 0\n        masked = difference.masked_fill(mask, 0)\n        return torch.mean(torch.pow(-masked, p))\n\n```",
    "937635": "Thanks for sharing!\nAnd Congratulations 2nd place!\n\nYou gave a lot of useful tips even at the beginning of the competition.\nThanks again! 👍",
    "937642": "Oh, thanks. It`s powerful. I learned a lot!",
    "937690": "Hello there,\n\nThanks for the quite detailed description and kudos for your second place ! :)\n\nQuick question however, you wrote \n*\"I trained my models on RGB input (using cv2.imread) and then fine-tuned **on non-rounded RGB images** (by manually decoding DCT-&gt;YCbCr -&gt; **DCT omitting rounding step**\"*\nDo you YCbCr --&gt; RGB ? or  fine-tuned on DCT (it would case it would be non-sense starting on RGB and fine-tuning on somethg completely different)\n\n\nBy the way, why didn' you use YCbCr directly (the data are indeed hidden in those channels) ?\nAnd how did you made this process ? I provided a code to get YCbCr data but kagglers complainted it is way too slow ? Using cv2 built-in functions ? \n\nOther questions, did you blended al QF together ?\n\nGreat post ... and great achievement ;)",
    "937705": "Hi @bloodaxe congrats\n\nI started this competition by reading your discussions in the [forum ](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/155392)on how to get started. Waiting for your kernel 💯",
    "937713": "Congratulations on your second place!\nThe tips you posted on the discussion were very informative and helpful.\nI have one question: what are the settings for optimizer and lr? \nIn my case I got CV 0.930 LB 0.927 with a multiclassification of B5 + adamW + ReduceLROnPlateau(init_lr=1e-3, factor=0.4, patience=2). I'm surprised you can get 0.935 LB with ResNet and DenseNet.",
    "937747": "Well summarized @bloodaxe It was a pleasure working with you! \n### A note on how to train SRNet\nA few competitors have experienced difficulties in training SRNet, which is expectable since it's training from scratch (and not from a well trained ImageNet model), some finesse is needed to make it converge. \nWe first trained on QF75 with the largest BS we could afford (BS=64) following the training schedule described in the SRNet paper, then we fine-tuned to the other QFs. This was done because (if we assume fixed payload) QF75 is usually easier to detect. In this case, even if the payload widely varied, QF75 and QF90 were still easier than QF95.\n### Don't use cv2.imread\nUnless you are willing to give up on free performance gains ;) We used a custom JPEG decoder which doesn't round/clip to [0,255], the gains can go up to 1% in wAUC (especially against JUNI). The forum had some examples of such decoders, we used our own implementation which takes advantage of numpy's stride_tricks for speed-ups (and jpegio of I/O). Some nets were first trained with cv2.imread then fine-tuned with the custom JPEG decoder.\nSRNet was trained on YCbCr, but the ImageNet pretrained models were trained on RGB (both non rounded). We believe that since ImageNet models were trained on RGB data, this would be more suitable. But in the end, it's a pretty easy linear transformation, I wouldn't be surprised if YCbCr worked as well.\n### A comparison of some of our models\nBesides SRNet, the rest of the nets were trained on all QFs without any architecture modification. Note that the scatter shifts up when replacing Swish with Mish activation (c.f. original post), we couldn't do this for all our models in time ...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1012681%2Fcf6a6c7f0c49fd6c34d158048b501084%2Fall_stego.png?generation=1595312025650266&amp;alt=media)\n### Good old hand crafted features\nWe stuck in some good old steganalysis hand crafted features (DCTR and JRM + FLD ensemble, which can both be found [on our website](http://dde.binghamton.edu/download/feature_extractors/)), they didn't perform the best (wAUC=0.850) but were added for the sake of diversity.\n### The stacking\nWe stacked our models and trained Catboost on holdout set (all QFs together), the QF was added as a categorical feature. Hyper parameters were tuned using [SKOPT BayesSearchCV](https://scikit-optimize.github.io/stable/modules/generated/skopt.BayesSearchCV.html), but I'm sure a large enough grid search can do the job. Eugene Had a similar approach using xgboost.\n\n**Thanks to @remicogranne , Patrick, Quentin and the Kaggle team for this competition! And thanks to everyone who made this competition so much fun!**\n*More on our coming IEEE WIFS submission.*",
    "937857": "Thanks for sharing and a good explanation! 👍",
    "937877": "Thank you! \nI experimented with many hyper-parameters, and for me following combination worked quite good:\n- `SGD (lr=1e-2, wd=1e-4/1e-5)`, with cosine annealing to 1e-5.\n- `RAdam` (If starting from ImageNet) and `AdamW` (fine-tuning) with `lr=1e-4` and `wd=1e-2`.\n- `Ranger` with `flat + cos` schedule (For first half of the epochs LR is fixed, then decay with cosine to 0.01 of initial LR).\nB6/B7 models had dropout 0.5 after global average pooling.",
    "937886": "Pardon for the typo in the summary. I tried to write asap and made a mistake in the description. \nThe conversion was DCT -&gt; YCbCr -&gt; RGB. But we computed non-rounded float32 RGB values. They were quite close to values one may get from `cv2.imread`, but this non-rounded input had an additional signal, which improved performance of every model we fine-tuned on non-rounded image input. Hope it clarifies our data pipeline.\n\nAs for YCbCr, I initially experimented with it, but the CV/LB score was quite low, so I put on ice this approach in favor of RGB (known to work well) and test as many architectures I can. Same was with DCT input. I have almost identical approach as in 1st place solution, but haven't figured out how to incorporate embedding from DCT models in the second-level stacking. \n\nI included QF only in the second-level stacking model as one-hot vector. CV score with or without this feature was the same (maybe the difference was somewhere in 5 or 6 digit) but I decided to keep it just in case.",
    "937889": "Thank you for the detailed explanation!\nDid you plan to explain these parts later in the paper? 😊 \n\nAnyway, Congratulation 2nd place!",
    "937972": "Congrats @bloodaxe Could you please share code. It always excellent learning going through your elegantly written code",
    "938102": "\"Do you YCbCr --&gt; RGB ? or fine-tuned on DCT (it would case it would be non-sense starting on RGB and fine-tuning on somethg completely different)\"\n\npretrain on RGB and later retrain on YCbCr actually works. I start with RGB and was unable to get good results. then i switch to YCbCr. to save time, i reuse rgb models as initiation when training ycbcr.\n\nI guess it can work because the gray channel in ycbcr is close to rgb, so the network just need to learn two more channels cb and cr. but the information in these two channels are not that strong",
    "938671": "Thanks for sharing your experince in detail @bloodaxe it is very illuminating!"
  },
  "source": "meta"
}