{
  "id": 148162,
  "title": "Best single model",
  "url": "/competitions/alaska2-image-steganalysis/discussion/148162",
  "author_name": "",
  "post_date": "2020-05-03T11:16:54.065366500Z",
  "votes": 39,
  "comment_count": 97,
  "views": 0,
  "content": "<p>Let's share some insight. =)</p>\n\n<p><strong>Baseline</strong> (w/o full data set)\n<code>\nModel: EfficientNet\nEpoch: 5\nCV: 0.71\nLB: 0.75\n</code></p>\n\n<h2>Update</h2>\n\n<p>nearly full data set.\n<code>\nModel: EfficientNet\nEpoch: ~15\nCV: 0.88\nLB: 0.84\n</code></p>\n\n<h2>Update</h2>\n\n<p><code>\nModel: MixNet\nCV: 0.913\nLB: 0.910\n</code></p>",
  "messages": [
    {
      "id": "831429",
      "postDate": "05/03/2020 11:16:54",
      "content": "<p>Let's share some insight. =)</p>\n\n<p><strong>Baseline</strong> (w/o full data set)\n<code>\nModel: EfficientNet\nEpoch: 5\nCV: 0.71\nLB: 0.75\n</code></p>\n\n<h2>Update</h2>\n\n<p>nearly full data set.\n<code>\nModel: EfficientNet\nEpoch: ~15\nCV: 0.88\nLB: 0.84\n</code></p>\n\n<h2>Update</h2>\n\n<p><code>\nModel: MixNet\nCV: 0.913\nLB: 0.910\n</code></p>",
      "rawMarkdown": "Let's share some insight. =)\n\n**Baseline** (w/o full data set)\n```\nModel: EfficientNet\nEpoch: 5\nCV: 0.71\nLB: 0.75\n```\n\n## Update\n\nnearly full data set.\n```\nModel: EfficientNet\nEpoch: ~15\nCV: 0.88\nLB: 0.84\n```\n\n## Update\n\n```\nModel: MixNet\nCV: 0.913\nLB: 0.910\n```",
      "votes": null
    },
    {
      "id": "832371",
      "postDate": "05/04/2020 04:47:20",
      "content": "<p>```\nModel : EfficientNet-b3 (w/ some augs)\nResolution : 512x512x3 (rgb)\nEpoch : 12 (w/o EarlyStop)\nCV : 0.62 (90/10 split)\nLB : 0.801</p>\n\n<p>w/o full dataset\nw/o TTA\n```\nLB fluctuates depending on the seed. (it's not good)</p>",
      "rawMarkdown": "```\nModel : EfficientNet-b3 (w/ some augs)\nResolution : 512x512x3 (rgb)\nEpoch : 12 (w/o EarlyStop)\nCV : 0.62 (90/10 split)\nLB : 0.801\n\nw/o full dataset\nw/o TTA\n```\nLB fluctuates depending on the seed. (it's not good)",
      "votes": null
    },
    {
      "id": "832453",
      "postDate": "05/04/2020 06:31:33",
      "content": "<p>Are you training on TPU or GPU?</p>",
      "rawMarkdown": "Are you training on TPU or GPU?",
      "votes": null
    },
    {
      "id": "832485",
      "postDate": "05/04/2020 07:06:28",
      "content": "<p>training on TPU! \ncurrently, I'm trying to use higher resolution &amp; batch size, GPU can't handle it cuz of OOM :( so I use TPU.</p>",
      "rawMarkdown": "training on TPU! \ncurrently, I'm trying to use higher resolution &amp; batch size, GPU can't handle it cuz of OOM :( so I use TPU.",
      "votes": null
    },
    {
      "id": "832491",
      "postDate": "05/04/2020 07:15:56",
      "content": "<p>I am also facing the seed issue. However, I am also working on TPU but I think, the usage time limitation is a big barrier, on top that some unexpected issue related to TPU causes time kills :(</p>",
      "rawMarkdown": "I am also facing the seed issue. However, I am also working on TPU but I think, the usage time limitation is a big barrier, on top that some unexpected issue related to TPU causes time kills :(",
      "votes": null
    },
    {
      "id": "832851",
      "postDate": "05/04/2020 13:19:31",
      "content": "<p>I 'm also a.l.s.o. facing the seed issue. :(</p>\n\n<p>I already used TPU quota in this week.</p>",
      "rawMarkdown": "I 'm also a.l.s.o. facing the seed issue. :(\n\nI already used TPU quota in this week.",
      "votes": null
    },
    {
      "id": "834033",
      "postDate": "05/05/2020 08:39:47",
      "content": "<p>efficientnet-b0, 512, 512, 3\nlocal val: 0.890, public LB: 0.895</p>",
      "rawMarkdown": "efficientnet-b0, 512, 512, 3\nlocal val: 0.890, public LB: 0.895",
      "votes": null
    },
    {
      "id": "834039",
      "postDate": "05/05/2020 08:43:23",
      "content": "<p>Your CV and LB correlated. Have you tried on the whole dataset?</p>",
      "rawMarkdown": "Your CV and LB correlated. Have you tried on the whole dataset?",
      "votes": null
    },
    {
      "id": "834189",
      "postDate": "05/05/2020 11:23:35",
      "content": "<p>It's worth noting that the public LB is really small, only 1000 samples, so some fluctuation +-1% at least t is expected.</p>\n\n<p>I wonder why the organisers didn't make the test set bigger given they were able to obtain 75,000 images for training (different sources?)</p>",
      "rawMarkdown": "It's worth noting that the public LB is really small, only 1000 samples, so some fluctuation +-1% at least t is expected.\n\nI wonder why the organisers didn't make the test set bigger given they were able to obtain 75,000 images for training (different sources?)",
      "votes": null
    },
    {
      "id": "834190",
      "postDate": "05/05/2020 11:25:55",
      "content": "<p>Wow, that's a lot better than I got with efficientnet (0.853 CV, 0.880 LB)</p>",
      "rawMarkdown": "Wow, that's a lot better than I got with efficientnet (0.853 CV, 0.880 LB)",
      "votes": null
    },
    {
      "id": "834316",
      "postDate": "05/05/2020 13:35:55",
      "content": "<p>Yes, actually I split the validation set first, and gradually increase the number of training samples. During this process, in most of the time cv and lb are correlated.</p>",
      "rawMarkdown": "Yes, actually I split the validation set first, and gradually increase the number of training samples. During this process, in most of the time cv and lb are correlated.",
      "votes": null
    },
    {
      "id": "834320",
      "postDate": "05/05/2020 13:38:47",
      "content": "<p>I also got larger gap during my old experiments, FYI:\nlocal 0.845, LB 0.869\nlocal 0.860, LB 0.884</p>",
      "rawMarkdown": "I also got larger gap during my old experiments, FYI:\nlocal 0.845, LB 0.869\nlocal 0.860, LB 0.884",
      "votes": null
    },
    {
      "id": "836417",
      "postDate": "05/07/2020 01:07:06",
      "content": "<p>```\nArch : EfficientNet-b0\nResolution : 512x512x3\nEpoch : maybe 5 ~ 6 (took about 45 mins per epoch on V100 x 1)\nCV : 0.71 (acc) / 0.81 (weighted auc)\nLB : 0.844</p>\n\n<p>w/ whole dataset\nw/o TTA\n```</p>\n\n<p>------- updated</p>\n\n<p>```\nArch : EfficientNet\nCV : 0.83 (weighted auc)\nLB : 0.896</p>\n\n<p>w/ whole dataset\nw/ TTA\n```</p>\n\n<p>In my case, TTA helps a lot.</p>",
      "rawMarkdown": "```\nArch : EfficientNet-b0\nResolution : 512x512x3\nEpoch : maybe 5 ~ 6 (took about 45 mins per epoch on V100 x 1)\nCV : 0.71 (acc) / 0.81 (weighted auc)\nLB : 0.844\n\nw/ whole dataset\nw/o TTA\n```\n\n------- updated\n\n```\nArch : EfficientNet\nCV : 0.83 (weighted auc)\nLB : 0.896\n\nw/ whole dataset\nw/ TTA\n```\n\nIn my case, TTA helps a lot.",
      "votes": null
    },
    {
      "id": "836710",
      "postDate": "05/07/2020 07:06:18",
      "content": "<p>Did you use the whole dataset or just a subset? </p>",
      "rawMarkdown": "Did you use the whole dataset or just a subset?",
      "votes": null
    },
    {
      "id": "836747",
      "postDate": "05/07/2020 07:56:06",
      "content": "<p>Hmm, the gap between those two are significant. \nHow many time takes to complete on epoch in GPU? (I guess you're using the whole set).</p>",
      "rawMarkdown": "Hmm, the gap between those two are significant. \nHow many time takes to complete on epoch in GPU? (I guess you're using the whole set).",
      "votes": null
    },
    {
      "id": "836785",
      "postDate": "05/07/2020 08:28:03",
      "content": "<p>Model : EfficientNet\nLocal CV: 0.853  whole dataset (0.8:0.2)\nPublic LB: 0.889 \nPublic LB(TTA): 0.893</p>",
      "rawMarkdown": "Model : EfficientNet\nLocal CV: 0.853  whole dataset (0.8:0.2)\nPublic LB: 0.889 \nPublic LB(TTA): 0.893",
      "votes": null
    },
    {
      "id": "836814",
      "postDate": "05/07/2020 09:02:41",
      "content": "<p>that's great. are you training on GPU or TPU?</p>",
      "rawMarkdown": "that's great. are you training on GPU or TPU?",
      "votes": null
    },
    {
      "id": "836816",
      "postDate": "05/07/2020 09:04:05",
      "content": "<p>yeap! this time, i use the whole data set. it takes about 45 mins per epoch (V100 x 1).\nand i just do early stop (usually stopped at 5 ~ 6 epochs),</p>",
      "rawMarkdown": "yeap! this time, i use the whole data set. it takes about 45 mins per epoch (V100 x 1).\nand i just do early stop (usually stopped at 5 ~ 6 epochs),",
      "votes": null
    },
    {
      "id": "838298",
      "postDate": "05/08/2020 13:04:23",
      "content": "<p>Training on TPU, because the dataset is a little bit large. </p>\n\n<p>Update:\nLocal CV: 0.877\nPublic LB(TTA): 0.893</p>",
      "rawMarkdown": "Training on TPU, because the dataset is a little bit large. \n\nUpdate:\nLocal CV: 0.877\nPublic LB(TTA): 0.893",
      "votes": null
    },
    {
      "id": "841804",
      "postDate": "05/11/2020 03:08:41",
      "content": "<p>I have two models with the same split.\n1)\nLocal CV: 0.877\nPublic LB(TTA): 0.893\n2)\nLocal CV: 0.902\nPublic LB(TTA): 0.890</p>\n\n<p>Maybe +-2%  is expected.</p>",
      "rawMarkdown": "I have two models with the same split.\n1)\nLocal CV: 0.877\nPublic LB(TTA): 0.893\n2)\nLocal CV: 0.902\nPublic LB(TTA): 0.890\n\nMaybe +-2%  is expected.",
      "votes": null
    },
    {
      "id": "842363",
      "postDate": "05/11/2020 11:05:52",
      "content": "<p>Well done !\nI strongly invite you to test the following  (see my post <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/148919\">here</a> and <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/147494\">there</a> ):\n- Try multiclass for idenfying each embedding scheme individually\n- Try split the dataset in terms of JPEG quality factor (there is only 3 of them, you can easily guess by look at quantization table)\n- Increase the training set (possibly using data augmentation , <a href=\"https://hal-lirmm.ccsd.cnrs.fr/lirmm-02559838/file/IHMMSec-2016_Yedroudj_Chaumont_Comby_Amara_Bas_Pixels-off.pdf\">see also this paper that try many possible ways and show there efficiency</a>)\n- Try incorporating DCT coefficients\n- and, of course, other tricks that you guys know much better than me such as using several models and merge them (as already proposed by merging efficient net b3 and b7 ...) </p>",
      "rawMarkdown": "Well done !\nI strongly invite you to test the following  (see my post [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/148919) and [there](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/147494) ):\n- Try multiclass for idenfying each embedding scheme individually\n- Try split the dataset in terms of JPEG quality factor (there is only 3 of them, you can easily guess by look at quantization table)\n- Increase the training set (possibly using data augmentation , [see also this paper that try many possible ways and show there efficiency](https://hal-lirmm.ccsd.cnrs.fr/lirmm-02559838/file/IHMMSec-2016_Yedroudj_Chaumont_Comby_Amara_Bas_Pixels-off.pdf))\n- Try incorporating DCT coefficients\n- and, of course, other tricks that you guys know much better than me such as using several models and merge them (as already proposed by merging efficient net b3 and b7 ...)",
      "votes": null
    },
    {
      "id": "842370",
      "postDate": "05/11/2020 11:08:00",
      "content": "<p>Yup ! There may be a small difference because of the small testing set.\nTo be clear, the generation of the dataset is extremely time consuming and, in this context, getting a very large training dataset may be important if you want to consider a targeted attack (by spliting the data by JPEG quality factors, embedding schemes, noise level, processing operations etc ...\nWe have noted that this can clearly help a lot ... yet this requires more training samples </p>",
      "rawMarkdown": "Yup ! There may be a small difference because of the small testing set.\nTo be clear, the generation of the dataset is extremely time consuming and, in this context, getting a very large training dataset may be important if you want to consider a targeted attack (by spliting the data by JPEG quality factors, embedding schemes, noise level, processing operations etc ...\nWe have noted that this can clearly help a lot ... yet this requires more training samples",
      "votes": null
    },
    {
      "id": "842377",
      "postDate": "05/11/2020 11:11:17",
      "content": "<p>Here's a snippet to split by JPEG quality factor (it is a public data, it is simply in the header of the JPEG files ....)\n<code>for file in train_files[start: end]:</code>\n<code>____fullCover = folders[3] + '/' + file</code>\n<code>____randFolder = np.random.randint(0,3)</code>\n<code>____full = folders[randFolder] + '/' + file</code>\n<code>____imgJPEG = jio.read(path+folders[0]+'/'+file)</code>\n<code>____if( imgJPEG.quant_tables[0][0,0] == 2):</code>\n<code>________cover_paths95.append(fullCover)</code>\n<code>________steg_paths95.append(full)</code>\n<code>____elif( imgJPEG.quant_tables[0][0,0] == 3):</code>\n<code>________cover_paths90.append(fullCover)</code>\n<code>________steg_paths90.append(full)</code>\n<code>____elif( imgJPEG.quant_tables[0][0,0] == 8):</code>\n<code>________cover_paths75.append(fullCover)</code>\n<code>________steg_paths75.append(full)</code></p>",
      "rawMarkdown": "Here's a snippet to split by JPEG quality factor (it is a public data, it is simply in the header of the JPEG files ....)\n`for file in train_files[start: end]:`\n`____fullCover = folders[3] + '/' + file`\n`____randFolder = np.random.randint(0,3)`\n`____full = folders[randFolder] + '/' + file`\n`____imgJPEG = jio.read(path+folders[0]+'/'+file)`\n`____if( imgJPEG.quant_tables[0][0,0] == 2):`\n`________cover_paths95.append(fullCover)`\n`________steg_paths95.append(full)`\n`____elif( imgJPEG.quant_tables[0][0,0] == 3):`\n`________cover_paths90.append(fullCover)`\n`________steg_paths90.append(full)`\n`____elif( imgJPEG.quant_tables[0][0,0] == 8):`\n`________cover_paths75.append(fullCover)`\n`________steg_paths75.append(full)`",
      "votes": null
    },
    {
      "id": "842386",
      "postDate": "05/11/2020 11:16:11",
      "content": "<p><a href=\"/remicogranne\">@remicogranne</a> thanks, it helps indeed. :)</p>",
      "rawMarkdown": "remicogranne thanks, it helps indeed. :)",
      "votes": null
    },
    {
      "id": "842565",
      "postDate": "05/11/2020 13:22:46",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "844675",
      "postDate": "05/12/2020 19:05:37",
      "content": "<p>Thank you for sharing! Can I ask you a question, did you do it as binary classification or multi? </p>",
      "rawMarkdown": "Thank you for sharing! Can I ask you a question, did you do it as binary classification or multi?",
      "votes": null
    },
    {
      "id": "844898",
      "postDate": "05/12/2020 23:10:32",
      "content": "<p>Check <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/150359\">here</a></p>",
      "rawMarkdown": "Check [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/150359)",
      "votes": null
    },
    {
      "id": "846720",
      "postDate": "05/14/2020 00:30:46",
      "content": "<p>MixNet\nCV: 0.889 (90/10 split)\nLB: 0.901 \nLB (TTA): 0.894</p>\n\n<p>Not sure why TTA worsens score. I'm doing hflip, vflip, and transpose height/width. </p>",
      "rawMarkdown": "MixNet\nCV: 0.889 (90/10 split)\nLB: 0.901 \nLB (TTA): 0.894\n\nNot sure why TTA worsens score. I'm doing hflip, vflip, and transpose height/width.",
      "votes": null
    },
    {
      "id": "852195",
      "postDate": "05/18/2020 08:22:07",
      "content": "<p>me too, have you figured out the reason?</p>",
      "rawMarkdown": "me too, have you figured out the reason?",
      "votes": null
    },
    {
      "id": "852225",
      "postDate": "05/18/2020 08:48:35",
      "content": "<p>Hi  <a href=\"/kozistr\">@kozistr</a>, what's the transformations in your TTA？</p>",
      "rawMarkdown": "Hi  @kozistr, what's the transformations in your TTA？",
      "votes": null
    },
    {
      "id": "852381",
      "postDate": "05/18/2020 11:46:26",
      "content": "<p>Vertical &amp; Horizontal flips are used!</p>",
      "rawMarkdown": "Vertical &amp; Horizontal flips are used!",
      "votes": null
    },
    {
      "id": "856951",
      "postDate": "05/22/2020 07:09:51",
      "content": "<p>efficientnet-b2 80/20 split\ncv(no tta): 913\nlb(with horizon/vertical flip TTA): 920</p>",
      "rawMarkdown": "efficientnet-b2 80/20 split\ncv(no tta): 913\nlb(with horizon/vertical flip TTA): 920",
      "votes": null
    },
    {
      "id": "857695",
      "postDate": "05/22/2020 20:40:12",
      "content": "<p>it makes sense to look on validation only in this competition) it is ok you get +-1% on LB</p>",
      "rawMarkdown": "it makes sense to look on validation only in this competition) it is ok you get +-1% on LB",
      "votes": null
    },
    {
      "id": "866749",
      "postDate": "05/29/2020 16:38:37",
      "content": "<p>Are you using DCT images? If you augment them, I think you might need to correct each 8x8 block so that it is orientated correctly</p>",
      "rawMarkdown": "Are you using DCT images? If you augment them, I think you might need to correct each 8x8 block so that it is orientated correctly",
      "votes": null
    },
    {
      "id": "867110",
      "postDate": "05/30/2020 02:33:19",
      "content": "<p>I tried it, but it's too slow.</p>",
      "rawMarkdown": "I tried it, but it's too slow.",
      "votes": null
    },
    {
      "id": "872588",
      "postDate": "06/03/2020 10:32:07",
      "content": "<p>Hello! 😁 And what is the \"jio\" library? I can't find it, so I am blindly assuming is a shortcut from an actual name... </p>",
      "rawMarkdown": "Hello! 😁 And what is the \"jio\" library? I can't find it, so I am blindly assuming is a shortcut from an actual name...",
      "votes": null
    },
    {
      "id": "873234",
      "postDate": "06/04/2020 00:10:20",
      "content": "<p><a href=\"/andradaolteanu\">@andradaolteanu</a> <code>jpegio</code></p>",
      "rawMarkdown": "andradaolteanu `jpegio`",
      "votes": null
    },
    {
      "id": "883756",
      "postDate": "06/12/2020 22:12:32",
      "content": "<p>So how much time does it take the model to train on the full data set for one epoch? For me it took about 2 hours on the Kaggle GPU but I want to know if that is a reasonable number.</p>",
      "rawMarkdown": "So how much time does it take the model to train on the full data set for one epoch? For me it took about 2 hours on the Kaggle GPU but I want to know if that is a reasonable number.",
      "votes": null
    },
    {
      "id": "883788",
      "postDate": "06/12/2020 23:37:30",
      "content": "<p>depends on your model architecture. For me it is about one hour on full dataset for eb0. </p>",
      "rawMarkdown": "depends on your model architecture. For me it is about one hour on full dataset for eb0.",
      "votes": null
    },
    {
      "id": "901705",
      "postDate": "06/25/2020 16:33:41",
      "content": "<p>efficientnet-b3 80/20 split (hold out)\ncv (w/o tta): 0.926\nlb (w/o tta): 0.925\nlb (with tta): 0.931</p>\n\n<p>How can we reach 0.940+...?</p>",
      "rawMarkdown": "efficientnet-b3 80/20 split (hold out)\ncv (w/o tta): 0.926\nlb (w/o tta): 0.925\nlb (with tta): 0.931\n\nHow can we reach 0.940+...?",
      "votes": null
    },
    {
      "id": "901714",
      "postDate": "06/25/2020 16:37:26",
      "content": "<p>i think blending is the answer:) What is your cv with tta? I have cv with single model w/o tta 0.93, with tta 0.9335, but still lb was 0.93:(</p>",
      "rawMarkdown": "i think blending is the answer:) What is your cv with tta? I have cv with single model w/o tta 0.93, with tta 0.9335, but still lb was 0.93:(",
      "votes": null
    },
    {
      "id": "901749",
      "postDate": "06/25/2020 17:08:22",
      "content": "<p>blending... =(</p>\n\n<p>I tested cv with tta;\ncv (with tta): 0.930</p>",
      "rawMarkdown": "blending... =(\n\nI tested cv with tta;\ncv (with tta): 0.930",
      "votes": null
    },
    {
      "id": "902053",
      "postDate": "06/25/2020 22:23:28",
      "content": "<p>My ef-b3 also has a good val score but not high score on LB, do you noticed similar thing? Also by cv do you use 5fold model ?</p>",
      "rawMarkdown": "My ef-b3 also has a good val score but not high score on LB, do you noticed similar thing? Also by cv do you use 5fold model ?",
      "votes": null
    },
    {
      "id": "902073",
      "postDate": "06/25/2020 23:00:00",
      "content": "<p>Nope, most representetive split of fixed images</p>",
      "rawMarkdown": "Nope, most representetive split of fixed images",
      "votes": null
    },
    {
      "id": "906126",
      "postDate": "06/29/2020 04:43:33",
      "content": "<p><a href=\"/kozistr\">@kozistr</a> Hi! Are you still using TPU? If so, are you using full dataset and full resolution? How did you bypass the memory issue.</p>",
      "rawMarkdown": "kozistr Hi! Are you still using TPU? If so, are you using full dataset and full resolution? How did you bypass the memory issue.",
      "votes": null
    },
    {
      "id": "909500",
      "postDate": "06/30/2020 16:12:58",
      "content": "<p>Hi <a href=\"/tonychenxyz\">@tonychenxyz</a> !\nNope, i'm currently using GPU w/ full dataset, resolution. \nYou can read the image files at runtime instead of loading full dataset on the memory at once.\nIt can solve the memory issue, but I/O read will cost. (but, in this case, it's not a big issue though :) )</p>",
      "rawMarkdown": "Hi @tonychenxyz !\nNope, i'm currently using GPU w/ full dataset, resolution. \nYou can read the image files at runtime instead of loading full dataset on the memory at once.\nIt can solve the memory issue, but I/O read will cost. (but, in this case, it's not a big issue though :) )",
      "votes": null
    },
    {
      "id": "910273",
      "postDate": "07/01/2020 04:47:00",
      "content": "<p>What strategies did TTA use? Instead, my LB(with TTA) dropped points</p>",
      "rawMarkdown": "What strategies did TTA use? Instead, my LB(with TTA) dropped points",
      "votes": null
    },
    {
      "id": "910661",
      "postDate": "07/01/2020 09:31:19",
      "content": "<p>great to see your results <a href=\"/ren4yu\">@ren4yu</a>  how are training on whole dataset is it on kaggle or on other platform thanks</p>",
      "rawMarkdown": "great to see your results @ren4yu  how are training on whole dataset is it on kaggle or on other platform thanks",
      "votes": null
    },
    {
      "id": "915965",
      "postDate": "07/05/2020 08:45:08",
      "content": "<p>Just started:</p>\n\n<p><code>\nefficientnet-b0\nCV: 0.916\nLB: 0.914\n</code></p>",
      "rawMarkdown": "Just started:\n\n```\nefficientnet-b0\nCV: 0.916\nLB: 0.914\n```",
      "votes": null
    },
    {
      "id": "916041",
      "postDate": "07/05/2020 09:50:46",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Good to see you 😃 </p>",
      "rawMarkdown": "philippsinger Good to see you 😃",
      "votes": null
    },
    {
      "id": "916846",
      "postDate": "07/06/2020 04:43:17",
      "content": "<p>efficientnet-b2</p>\n\n<p>cv 0.923\nlb 0.926 with tta,</p>\n\n<p>And does the big model better?</p>",
      "rawMarkdown": "efficientnet-b2\n\ncv 0.923\nlb 0.926 with tta,\n\nAnd does the big model better?",
      "votes": null
    },
    {
      "id": "917385",
      "postDate": "07/06/2020 13:10:16",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> \nAre you training on GPU or TPU? May I know how much time it took per epoch for e0?</p>",
      "rawMarkdown": "philippsinger \nAre you training on GPU or TPU? May I know how much time it took per epoch for e0?",
      "votes": null
    },
    {
      "id": "917388",
      "postDate": "07/06/2020 13:15:05",
      "content": "<p>have you calculted w/o tta? did it help? \nBig model is different, but training e5/6/7 took forever kind of :(</p>",
      "rawMarkdown": "have you calculted w/o tta? did it help? \nBig model is different, but training e5/6/7 took forever kind of :(",
      "votes": null
    },
    {
      "id": "917396",
      "postDate": "07/06/2020 13:23:13",
      "content": "<p>A single b1 model can achieve LB 0.931 w/ TTA.</p>",
      "rawMarkdown": "A single b1 model can achieve LB 0.931 w/ TTA.",
      "votes": null
    },
    {
      "id": "917452",
      "postDate": "07/06/2020 14:23:29",
      "content": "<p>GPU locally - it will always depend on your HW</p>",
      "rawMarkdown": "GPU locally - it will always depend on your HW",
      "votes": null
    },
    {
      "id": "917470",
      "postDate": "07/06/2020 14:36:30",
      "content": "<p><a href=\"/ren4yu\">@ren4yu</a> Good to know - but I assume there are some tricks necessary for that :)</p>",
      "rawMarkdown": "ren4yu Good to know - but I assume there are some tricks necessary for that :)",
      "votes": null
    },
    {
      "id": "918006",
      "postDate": "07/06/2020 21:56:08",
      "content": "<p>Model: SRNet\nEpochs: ~25\nLB: 0.845</p>",
      "rawMarkdown": "Model: SRNet\nEpochs: ~25\nLB: 0.845",
      "votes": null
    },
    {
      "id": "918138",
      "postDate": "07/07/2020 02:51:33",
      "content": "<p>i haven't test without tta. </p>",
      "rawMarkdown": "i haven't test without tta.",
      "votes": null
    },
    {
      "id": "918139",
      "postDate": "07/07/2020 02:52:55",
      "content": "<p>Thanks for that. <a href=\"/ren4yu\">@ren4yu</a> \nAnd one more question, how many epochs do your model trained? </p>",
      "rawMarkdown": "Thanks for that. @ren4yu \nAnd one more question, how many epochs do your model trained?",
      "votes": null
    },
    {
      "id": "920981",
      "postDate": "07/09/2020 01:59:50",
      "content": "<p>efficientnetB2\nsingle fold\nCV: 0.9079 LB: 0.907 (no tta)</p>",
      "rawMarkdown": "efficientnetB2\nsingle fold\nCV: 0.9079 LB: 0.907 (no tta)",
      "votes": null
    },
    {
      "id": "920992",
      "postDate": "07/09/2020 02:08:57",
      "content": "<p>TTA improvement is about +0.004 (for horizonal and veritical flip)</p>",
      "rawMarkdown": "TTA improvement is about +0.004 (for horizonal and veritical flip)",
      "votes": null
    },
    {
      "id": "921182",
      "postDate": "07/09/2020 05:41:01",
      "content": "<p>You implement your own SRNet? I also tried SRNet but that only allows me fed 2 images one batch (Kaggle notebook). May I ask what's your batch_size?</p>",
      "rawMarkdown": "You implement your own SRNet? I also tried SRNet but that only allows me fed 2 images one batch (Kaggle notebook). May I ask what's your batch_size?",
      "votes": null
    },
    {
      "id": "922177",
      "postDate": "07/09/2020 21:16:16",
      "content": "<p><a href=\"/vaillant\">@vaillant</a> \nensembling two models with TTA gave worse results, but nice boost separately. Didn't figure out why 😯 </p>",
      "rawMarkdown": "vaillant \nensembling two models with TTA gave worse results, but nice boost separately. Didn't figure out why 😯",
      "votes": null
    },
    {
      "id": "922178",
      "postDate": "07/09/2020 21:18:35",
      "content": "<p><a href=\"/ren4yu\">@ren4yu</a> \nhow many epochs have you trained for E3? I've traind E5 pretty long but didn't get any decent score like you. </p>",
      "rawMarkdown": "ren4yu \nhow many epochs have you trained for E3? I've traind E5 pretty long but didn't get any decent score like you.",
      "votes": null
    },
    {
      "id": "922180",
      "postDate": "07/09/2020 21:22:27",
      "content": "<p>It's somewhat misleading. It's not about your model score, but with the same model some other reporting better score. So, I think it's not proper to say only the base model, like others it has to do something with the head of this base model.  </p>",
      "rawMarkdown": "It's somewhat misleading. It's not about your model score, but with the same model some other reporting better score. So, I think it's not proper to say only the base model, like others it has to do something with the head of this base model.",
      "votes": null
    },
    {
      "id": "922245",
      "postDate": "07/10/2020 00:38:36",
      "content": "<p>I think 150 epochs would be enough. Yes, it requires really long time...</p>",
      "rawMarkdown": "I think 150 epochs would be enough. Yes, it requires really long time...",
      "votes": null
    },
    {
      "id": "922250",
      "postDate": "07/10/2020 00:42:15",
      "content": "<p>I think 150 epochs would be enough. Yes, it requires really long time. I trained on my local GPU machine early in this competition when I was very busy, and I left my machine training a same model for a long time...</p>",
      "rawMarkdown": "I think 150 epochs would be enough. Yes, it requires really long time. I trained on my local GPU machine early in this competition when I was very busy, and I left my machine training a same model for a long time...",
      "votes": null
    },
    {
      "id": "922321",
      "postDate": "07/10/2020 02:57:41",
      "content": "<p>LOL,  it is really a lone time. </p>",
      "rawMarkdown": "LOL,  it is really a lone time.",
      "votes": null
    },
    {
      "id": "924558",
      "postDate": "07/11/2020 13:58:05",
      "content": "<p>150 epoch is too long. \n<a href=\"/cooolz\">@cooolz</a> how many epochs have you trained for this?</p>",
      "rawMarkdown": "150 epoch is too long. \n@cooolz how many epochs have you trained for this?",
      "votes": null
    },
    {
      "id": "924581",
      "postDate": "07/11/2020 14:11:39",
      "content": "<p>🐱 </p>",
      "rawMarkdown": "🐱",
      "votes": null
    },
    {
      "id": "926132",
      "postDate": "07/12/2020 14:17:07",
      "content": "<p>Had originally experimented with effnet b0 based on forums recommendations. This was an excellent idea. It's super fast and reasonable. Started getting results that looked okay (they matched the results of the best public kernel, even though b0) so I got thirsty/greedy and made the huge mistake of starting to train a b3 model. Each epoch takes about an hour and some change on my dual 2080Ti machine. Since time is ticking, decided to make a sub off of current best epoch (31 out of 34 completed). From my observations, it is not overfit yet and can still maybe train an additional 15-20 epochs(?) to secure maybe an additional +0.005 I think. This is a single fold pytorch model trained on 80% of the data (GKF=5). I also have 3 other architectures I've designed but hadn't had compute resources to test up yet.</p>\n\n<p>B3 31 epoch\nTrain: CV 0.91857\nVal: CV 0.92084\nVal w/ TTA: 0.9289\nSub w/ TTA: 0.929</p>\n\n<p>B0 37 epoch (same architectural changes + but w/ less train augmentation and no TTA):\nTrain: CV 0.91948\nVal: CV 0.92011\nSub w/o TTA: 0.917</p>",
      "rawMarkdown": "Had originally experimented with effnet b0 based on forums recommendations. This was an excellent idea. It's super fast and reasonable. Started getting results that looked okay (they matched the results of the best public kernel, even though b0) so I got thirsty/greedy and made the huge mistake of starting to train a b3 model. Each epoch takes about an hour and some change on my dual 2080Ti machine. Since time is ticking, decided to make a sub off of current best epoch (31 out of 34 completed). From my observations, it is not overfit yet and can still maybe train an additional 15-20 epochs(?) to secure maybe an additional +0.005 I think. This is a single fold pytorch model trained on 80% of the data (GKF=5). I also have 3 other architectures I've designed but hadn't had compute resources to test up yet.\n\nB3 31 epoch\nTrain: CV 0.91857\nVal: CV 0.92084\nVal w/ TTA: 0.9289\nSub w/ TTA: 0.929\n\n\nB0 37 epoch (same architectural changes + but w/ less train augmentation and no TTA):\nTrain: CV 0.91948\nVal: CV 0.92011\nSub w/o TTA: 0.917",
      "votes": null
    },
    {
      "id": "926213",
      "postDate": "07/12/2020 15:21:59",
      "content": "<p>And how does B0 with TTA on CV/LB?</p>\n\n<p>CV 0.919 without TTA on B0 is pretty good, seems to match your B3 (even if trained a bit longer).</p>",
      "rawMarkdown": "And how does B0 with TTA on CV/LB?\n\nCV 0.919 without TTA on B0 is pretty good, seems to match your B3 (even if trained a bit longer).",
      "votes": null
    },
    {
      "id": "926244",
      "postDate": "07/12/2020 15:40:11",
      "content": "<blockquote>\n  <p>And how does B0 with TTA on CV/LB?</p>\n</blockquote>\n\n<p>No clue. I still have the checkpoints + notebook from the b0 model so I guess I could use kaggle kernel to run TTA on it(?).</p>\n\n<p>B3 model is still training locally + val CV just hit 0.92252 at epoch 35. Submitted and score is 0.931-LB. I think it might be time to figure out how this whole tpu thing works. My understanding is that the TPU node is hosted physically close to the computer doing training, ideally on the same subnet. The major bottleneck I'm seeing on my machine is decoding of JPEG. I tried caching the decoded jpeg's as .npy and as .pkl but then my SSD's light would come on and I'd hit transfer speed bottlenecks instead of decoding speed bottlenecks.</p>\n\n<p>Are you using TPU? Are you aware of the throughput of google's SSD? If they aren't remarkable, then I'd assume that the same transfer bottlenecks would occur, limiting the amount of data that could be pumped into the TPU........</p>",
      "rawMarkdown": "&gt; And how does B0 with TTA on CV/LB?\n\nNo clue. I still have the checkpoints + notebook from the b0 model so I guess I could use kaggle kernel to run TTA on it(?).\n\nB3 model is still training locally + val CV just hit 0.92252 at epoch 35. Submitted and score is 0.931-LB. I think it might be time to figure out how this whole tpu thing works. My understanding is that the TPU node is hosted physically close to the computer doing training, ideally on the same subnet. The major bottleneck I'm seeing on my machine is decoding of JPEG. I tried caching the decoded jpeg's as .npy and as .pkl but then my SSD's light would come on and I'd hit transfer speed bottlenecks instead of decoding speed bottlenecks.\n\nAre you using TPU? Are you aware of the throughput of google's SSD? If they aren't remarkable, then I'd assume that the same transfer bottlenecks would occur, limiting the amount of data that could be pumped into the TPU........",
      "votes": null
    },
    {
      "id": "932105",
      "postDate": "07/16/2020 17:43:42",
      "content": "<p>If you use pinned memory and multiple workers in your dataloader, JPG decoding shouldn't be a problem. Are you working on Linux?</p>",
      "rawMarkdown": "If you use pinned memory and multiple workers in your dataloader, JPG decoding shouldn't be a problem. Are you working on Linux?",
      "votes": null
    },
    {
      "id": "932160",
      "postDate": "07/16/2020 19:05:35",
      "content": "<p><a href=\"/authman\">@authman</a> TTA boosting 8 pts? Thats quite a bit...</p>",
      "rawMarkdown": "authman TTA boosting 8 pts? Thats quite a bit...",
      "votes": null
    },
    {
      "id": "932192",
      "postDate": "07/16/2020 19:49:50",
      "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> Was not using pinned memory. I'll give it a try, TY.</p>\n\n<p><a href=\"/philippsinger\">@philippsinger</a> ikr? TTA seems to work particularly well with EffNet. I've experimented with both B0 and B3. However, the exact same TTA with some of the other networks I've seen people use from the External Data Thread like PyConvNet, RexNet, ResNet, etc. etc. don't have as spectacular results. For example, yesterday's run:</p>\n\n<p>Epoch 42\nTrain: 0.92658\nVal: 0.92563\nVal TTA: 0.92819\nLB: 0.927</p>\n\n<p>I haven't started looking at model correlation yet. Will do so at the last moment when there's no more time to train.</p>",
      "rawMarkdown": "bigironsphere Was not using pinned memory. I'll give it a try, TY.\n\n@philippsinger ikr? TTA seems to work particularly well with EffNet. I've experimented with both B0 and B3. However, the exact same TTA with some of the other networks I've seen people use from the External Data Thread like PyConvNet, RexNet, ResNet, etc. etc. don't have as spectacular results. For example, yesterday's run:\n\nEpoch 42\nTrain: 0.92658\nVal: 0.92563\nVal TTA: 0.92819\nLB: 0.927\n\nI haven't started looking at model correlation yet. Will do so at the last moment when there's no more time to train.",
      "votes": null
    },
    {
      "id": "932196",
      "postDate": "07/16/2020 19:50:54",
      "content": "<p>0.930 -&gt; 0.926 (with TTA) for me, EfficientNet-B1</p>",
      "rawMarkdown": "0.930 -&gt; 0.926 (with TTA) for me, EfficientNet-B1",
      "votes": null
    },
    {
      "id": "932199",
      "postDate": "07/16/2020 19:57:30",
      "content": "<p><a href=\"/authman\">@authman</a> \nWhat's the result of your experiment on PyConvNet? Is it from this <a href=\"https://github.com/iduta/pyconv\">implementation</a>? Have you tried on GhostNet or Xception?</p>\n\n<p><a href=\"/vaillant\">@vaillant</a> \nTTA works well in my cases, just horizontal/vertical flip average. May I know what you've followed for this? I'm worried about TTA, will it overfit on private set!</p>",
      "rawMarkdown": "authman \nWhat's the result of your experiment on PyConvNet? Is it from this [implementation](https://github.com/iduta/pyconv)? Have you tried on GhostNet or Xception?\n\n@vaillant \nTTA works well in my cases, just horizontal/vertical flip average. May I know what you've followed for this? I'm worried about TTA, will it overfit on private set!",
      "votes": null
    },
    {
      "id": "932202",
      "postDate": "07/16/2020 20:03:40",
      "content": "<p>Fascinating. Perhaps the better base model you have, the less TTA works. We'll need some of the +top20's to weigh in.</p>\n\n<p>Yup, that was the PCN implementation I used.</p>",
      "rawMarkdown": "Fascinating. Perhaps the better base model you have, the less TTA works. We'll need some of the +top20's to weigh in.\n\nYup, that was the PCN implementation I used.",
      "votes": null
    },
    {
      "id": "932868",
      "postDate": "07/17/2020 11:07:01",
      "content": "<blockquote>\n  <p><strong>Ian Pan wrote:</strong></p>\n  \n  <p>0.930 -&gt; 0.926 (with TTA) for me, EfficientNet-B1</p>\n</blockquote>\n\n<p>our observation is that the relationship between CV/LB is fold-dependent.\nIn our 5 folds, we have some 3 folds get better CV than LB, the other two the other way around. and that is quite consistent both to different models, and among different iterations of the same model</p>\n\n<p>It might be a bit of a lottery in the end, because it is quite likely the private LB would yield relatively different score as compared to public LB</p>",
      "rawMarkdown": "&gt; **Ian Pan wrote:**\n&gt; \n&gt; 0.930 -&gt; 0.926 (with TTA) for me, EfficientNet-B1\n\nour observation is that the relationship between CV/LB is fold-dependent.\nIn our 5 folds, we have some 3 folds get better CV than LB, the other two the other way around. and that is quite consistent both to different models, and among different iterations of the same model\n\nIt might be a bit of a lottery in the end, because it is quite likely the private LB would yield relatively different score as compared to public LB",
      "votes": null
    },
    {
      "id": "932882",
      "postDate": "07/17/2020 11:14:19",
      "content": "<p><a href=\"/yifanxie\">@yifanxie</a> \nI have also observed similar stuff on fold-dependencies stuff. The ensemble of different models on the same fold with TTA improves significant but little on the different folds. </p>",
      "rawMarkdown": "yifanxie \nI have also observed similar stuff on fold-dependencies stuff. The ensemble of different models on the same fold with TTA improves significant but little on the different folds.",
      "votes": null
    },
    {
      "id": "932891",
      "postDate": "07/17/2020 11:17:46",
      "content": "<p>However, In our observation, small variants of model score much better than their larger variants with a reasonable config for both. </p>",
      "rawMarkdown": "However, In our observation, small variants of model score much better than their larger variants with a reasonable config for both.",
      "votes": null
    },
    {
      "id": "932902",
      "postDate": "07/17/2020 11:26:10",
      "content": "<blockquote>\n  <p><strong>M.Innat wrote:</strong></p>\n  \n  <p>However, In our observation, small variants of model score much better than their larger variants with a reasonable config for both. </p>\n</blockquote>\n\n<p>If I understand correctly, you are saying it is easier to get good score with smaller models (i.e. B0/B1) than larger models (i.e. B6/B7)?</p>\n\n<p>Our larger models are still in the oven, so we will see - our latest model is due to finished 2 hours before the deadline assuming no power outage 😌 </p>",
      "rawMarkdown": "&gt; **M.Innat wrote:**\n&gt; \n&gt; However, In our observation, small variants of model score much better than their larger variants with a reasonable config for both. \n\nIf I understand correctly, you are saying it is easier to get good score with smaller models (i.e. B0/B1) than larger models (i.e. B6/B7)?\n\nOur larger models are still in the oven, so we will see - our latest model is due to finished 2 hours before the deadline assuming no power outage 😌",
      "votes": null
    },
    {
      "id": "932927",
      "postDate": "07/17/2020 11:43:27",
      "content": "<p><a href=\"/yifanxie\">@yifanxie</a> \nKinda. Larger variants of EfficientNet (B7) didn't give a better score than its smaller variants. Also surprisingly lower than MixNet. As <a href=\"/ren4yu\">@ren4yu</a> <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/148162#917396\">mentioned earlier</a>, is kind of holds true here. But, it does not give any conclusion, if I'm not wrong, efficientnet is somewhat resolution-dependent. So, I was wondering if E7 needs a resolution &gt; than 512.</p>",
      "rawMarkdown": "yifanxie \nKinda. Larger variants of EfficientNet (B7) didn't give a better score than its smaller variants. Also surprisingly lower than MixNet. As @ren4yu [mentioned earlier](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/148162#917396), is kind of holds true here. But, it does not give any conclusion, if I'm not wrong, efficientnet is somewhat resolution-dependent. So, I was wondering if E7 needs a resolution &gt; than 512.",
      "votes": null
    },
    {
      "id": "932929",
      "postDate": "07/17/2020 11:43:51",
      "content": "<blockquote>\n  <p>our latest model is due to finished 2 hours before the deadline assuming no power outage</p>\n</blockquote>\n\n<p>Living on the edges 😎 </p>",
      "rawMarkdown": "&gt; our latest model is due to finished 2 hours before the deadline assuming no power outage\n\nLiving on the edges 😎",
      "votes": null
    },
    {
      "id": "932955",
      "postDate": "07/17/2020 11:56:21",
      "content": "<blockquote>\n  <p><strong>Ahmet Erdem wrote:</strong>\n  Living on the edges 😎 </p>\n</blockquote>\n\n<p>more like beggar can't be chooser 🙌 </p>",
      "rawMarkdown": "&gt; **Ahmet Erdem wrote:**\n&gt; Living on the edges 😎 \n\nmore like beggar can't be chooser 🙌",
      "votes": null
    },
    {
      "id": "932975",
      "postDate": "07/17/2020 12:02:55",
      "content": "<p>Is this competition more computationally expensive than DeepFake one btw?</p>",
      "rawMarkdown": "Is this competition more computationally expensive than DeepFake one btw?",
      "votes": null
    },
    {
      "id": "932992",
      "postDate": "07/17/2020 12:09:36",
      "content": "<p>From me, I wouldn't have any chances to compete with this <strong>computation hunger</strong> competition without my company support. Feel lucky to be a part of such corporations. </p>",
      "rawMarkdown": "From me, I wouldn't have any chances to compete with this **computation hunger** competition without my company support. Feel lucky to be a part of such corporations.",
      "votes": null
    },
    {
      "id": "932998",
      "postDate": "07/17/2020 12:14:17",
      "content": "<blockquote>\n  <p><strong>Ahmet Erdem wrote:</strong>\n  Is this competition more computationally expensive than DeepFake one btw?</p>\n</blockquote>\n\n<p>I am personally using 3~4x more GPU resources 😿 </p>",
      "rawMarkdown": "&gt; **Ahmet Erdem wrote:**\n&gt; Is this competition more computationally expensive than DeepFake one btw?\n\nI am personally using 3~4x more GPU resources 😿",
      "votes": null
    },
    {
      "id": "933055",
      "postDate": "07/17/2020 13:01:52",
      "content": "<p>Hopefully shakeup is kind to us. 😂 </p>\n\n<p>I was referring to my score on public LB with/without TTA. I'm using anokas' weighted AUC metric, and I had to quick fix a bug that would occur. However this bug gives me incorrect AUC values locally after a certain point (goes from 0.91/0.92-&gt;0.35), so not even sure what my CV is...</p>\n\n<p>For vanilla EfficientNet models, I found that larger models do better than smaller models. It will be interesting to see everyone's approaches at competition end.</p>",
      "rawMarkdown": "Hopefully shakeup is kind to us. 😂 \n\nI was referring to my score on public LB with/without TTA. I'm using anokas' weighted AUC metric, and I had to quick fix a bug that would occur. However this bug gives me incorrect AUC values locally after a certain point (goes from 0.91/0.92-&gt;0.35), so not even sure what my CV is...\n\nFor vanilla EfficientNet models, I found that larger models do better than smaller models. It will be interesting to see everyone's approaches at competition end.",
      "votes": null
    },
    {
      "id": "933069",
      "postDate": "07/17/2020 13:08:34",
      "content": "<p>I feel really curious to see the solution of the final top teams. And yes, the shakeup is kind to us 😄 </p>",
      "rawMarkdown": "I feel really curious to see the solution of the final top teams. And yes, the shakeup is kind to us 😄",
      "votes": null
    },
    {
      "id": "933510",
      "postDate": "07/17/2020 18:29:58",
      "content": "<p>Around 1 month ago, I got sth like: improve from local 9245 to 9294, but got LB 932 and 927 respectively. 😂 </p>",
      "rawMarkdown": "Around 1 month ago, I got sth like: improve from local 9245 to 9294, but got LB 932 and 927 respectively. 😂",
      "votes": null
    },
    {
      "id": "933515",
      "postDate": "07/17/2020 18:34:49",
      "content": "<p>That's long...😄 </p>",
      "rawMarkdown": "That's long...😄",
      "votes": null
    },
    {
      "id": "937375",
      "postDate": "07/21/2020 01:18:49",
      "content": "<p>B7 w/bitmix + TTA+ train on all data: 0.937 public/0.925 private</p>",
      "rawMarkdown": "B7 w/bitmix + TTA+ train on all data: 0.937 public/0.925 private",
      "votes": null
    },
    {
      "id": "937425",
      "postDate": "07/21/2020 02:33:57",
      "content": "<p>B7 + TTA + train on all data + finetune upper layers = 0.945 public / 0.926 private.</p>",
      "rawMarkdown": "B7 + TTA + train on all data + finetune upper layers = 0.945 public / 0.926 private.",
      "votes": null
    },
    {
      "id": "937446",
      "postDate": "07/21/2020 03:04:37",
      "content": "<p>B6 + TTA \nLocal / Public / Private: 0.940 / 0.940 / 0.929</p>",
      "rawMarkdown": "B6 + TTA \nLocal / Public / Private: 0.940 / 0.940 / 0.929",
      "votes": null
    },
    {
      "id": "937970",
      "postDate": "07/21/2020 09:22:28",
      "content": "<p>B4 Fold4 (public GKF notebook) local .931 private .925 (enough for 24th place)<br>\nB4 (5 folds) local .930 private .924 (enough for 30th place)</p>\n<p>Finally, I enjoyed learning about RegNetY, it was fun to train, and does so quite \"quickly\" (28epo@34min, single GPU)<br>\nRegNetY-8 (5 folds) +TTA local .924 private .920 (enough for 70th place)</p>",
      "rawMarkdown": "B4 Fold4 (public GKF notebook) local .931 private .925 (enough for 24th place)\nB4 (5 folds) local .930 private .924 (enough for 30th place)\n\nFinally, I enjoyed learning about RegNetY, it was fun to train, and does so quite \"quickly\" (28epo@34min, single GPU)\nRegNetY-8 (5 folds) +TTA local .924 private .920 (enough for 70th place)",
      "votes": null
    },
    {
      "id": "938020",
      "postDate": "07/21/2020 09:54:40",
      "content": "<p>effb4, single fold +  TTA (rot90, flips)\nlocal/public/private: 0.931/0.930/0.926</p>",
      "rawMarkdown": "effb4, single fold +  TTA (rot90, flips)\nlocal/public/private: 0.931/0.930/0.926",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 937970,
      "author_name": "robga",
      "author_url": "",
      "post_date": "07/21/2020 09:22:28",
      "content": "<p>B4 Fold4 (public GKF notebook) local .931 private .925 (enough for 24th place)<br>\nB4 (5 folds) local .930 private .924 (enough for 30th place)</p>\n<p>Finally, I enjoyed learning about RegNetY, it was fun to train, and does so quite \"quickly\" (28epo@34min, single GPU)<br>\nRegNetY-8 (5 folds) +TTA local .924 private .920 (enough for 70th place)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 832371,
      "author_name": "kozistr",
      "author_url": "",
      "post_date": "05/04/2020 04:47:20",
      "content": "<p>```\nModel : EfficientNet-b3 (w/ some augs)\nResolution : 512x512x3 (rgb)\nEpoch : 12 (w/o EarlyStop)\nCV : 0.62 (90/10 split)\nLB : 0.801</p>\n\n<p>w/o full dataset\nw/o TTA\n```\nLB fluctuates depending on the seed. (it's not good)</p>",
      "votes": null,
      "replies": [
        {
          "id": 832453,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "05/04/2020 06:31:33",
          "content": "<p>Are you training on TPU or GPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 832485,
          "author_name": "kozistr",
          "author_url": "",
          "post_date": "05/04/2020 07:06:28",
          "content": "<p>training on TPU! \ncurrently, I'm trying to use higher resolution &amp; batch size, GPU can't handle it cuz of OOM :( so I use TPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 832491,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "05/04/2020 07:15:56",
          "content": "<p>I am also facing the seed issue. However, I am also working on TPU but I think, the usage time limitation is a big barrier, on top that some unexpected issue related to TPU causes time kills :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 832851,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "05/04/2020 13:19:31",
          "content": "<p>I 'm also a.l.s.o. facing the seed issue. :(</p>\n\n<p>I already used TPU quota in this week.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 906126,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "06/29/2020 04:43:33",
          "content": "<p><a href=\"/kozistr\">@kozistr</a> Hi! Are you still using TPU? If so, are you using full dataset and full resolution? How did you bypass the memory issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 909500,
          "author_name": "kozistr",
          "author_url": "",
          "post_date": "06/30/2020 16:12:58",
          "content": "<p>Hi <a href=\"/tonychenxyz\">@tonychenxyz</a> !\nNope, i'm currently using GPU w/ full dataset, resolution. \nYou can read the image files at runtime instead of loading full dataset on the memory at once.\nIt can solve the memory issue, but I/O read will cost. (but, in this case, it's not a big issue though :) )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 834033,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "05/05/2020 08:39:47",
      "content": "<p>efficientnet-b0, 512, 512, 3\nlocal val: 0.890, public LB: 0.895</p>",
      "votes": null,
      "replies": [
        {
          "id": 834039,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "05/05/2020 08:43:23",
          "content": "<p>Your CV and LB correlated. Have you tried on the whole dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 834190,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "05/05/2020 11:25:55",
          "content": "<p>Wow, that's a lot better than I got with efficientnet (0.853 CV, 0.880 LB)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 834316,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "05/05/2020 13:35:55",
          "content": "<p>Yes, actually I split the validation set first, and gradually increase the number of training samples. During this process, in most of the time cv and lb are correlated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 834320,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "05/05/2020 13:38:47",
          "content": "<p>I also got larger gap during my old experiments, FYI:\nlocal 0.845, LB 0.869\nlocal 0.860, LB 0.884</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 842363,
          "author_name": "remicogranne",
          "author_url": "",
          "post_date": "05/11/2020 11:05:52",
          "content": "<p>Well done !\nI strongly invite you to test the following  (see my post <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/148919\">here</a> and <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/147494\">there</a> ):\n- Try multiclass for idenfying each embedding scheme individually\n- Try split the dataset in terms of JPEG quality factor (there is only 3 of them, you can easily guess by look at quantization table)\n- Increase the training set (possibly using data augmentation , <a href=\"https://hal-lirmm.ccsd.cnrs.fr/lirmm-02559838/file/IHMMSec-2016_Yedroudj_Chaumont_Comby_Amara_Bas_Pixels-off.pdf\">see also this paper that try many possible ways and show there efficiency</a>)\n- Try incorporating DCT coefficients\n- and, of course, other tricks that you guys know much better than me such as using several models and merge them (as already proposed by merging efficient net b3 and b7 ...) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 842377,
          "author_name": "remicogranne",
          "author_url": "",
          "post_date": "05/11/2020 11:11:17",
          "content": "<p>Here's a snippet to split by JPEG quality factor (it is a public data, it is simply in the header of the JPEG files ....)\n<code>for file in train_files[start: end]:</code>\n<code>____fullCover = folders[3] + '/' + file</code>\n<code>____randFolder = np.random.randint(0,3)</code>\n<code>____full = folders[randFolder] + '/' + file</code>\n<code>____imgJPEG = jio.read(path+folders[0]+'/'+file)</code>\n<code>____if( imgJPEG.quant_tables[0][0,0] == 2):</code>\n<code>________cover_paths95.append(fullCover)</code>\n<code>________steg_paths95.append(full)</code>\n<code>____elif( imgJPEG.quant_tables[0][0,0] == 3):</code>\n<code>________cover_paths90.append(fullCover)</code>\n<code>________steg_paths90.append(full)</code>\n<code>____elif( imgJPEG.quant_tables[0][0,0] == 8):</code>\n<code>________cover_paths75.append(fullCover)</code>\n<code>________steg_paths75.append(full)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 842386,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "05/11/2020 11:16:11",
          "content": "<p><a href=\"/remicogranne\">@remicogranne</a> thanks, it helps indeed. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 842565,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "05/11/2020 13:22:46",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 872588,
          "author_name": "andradaolteanu",
          "author_url": "",
          "post_date": "06/03/2020 10:32:07",
          "content": "<p>Hello! 😁 And what is the \"jio\" library? I can't find it, so I am blindly assuming is a shortcut from an actual name... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 873234,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "06/04/2020 00:10:20",
          "content": "<p><a href=\"/andradaolteanu\">@andradaolteanu</a> <code>jpegio</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 834189,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "05/05/2020 11:23:35",
      "content": "<p>It's worth noting that the public LB is really small, only 1000 samples, so some fluctuation +-1% at least t is expected.</p>\n\n<p>I wonder why the organisers didn't make the test set bigger given they were able to obtain 75,000 images for training (different sources?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 841804,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "05/11/2020 03:08:41",
          "content": "<p>I have two models with the same split.\n1)\nLocal CV: 0.877\nPublic LB(TTA): 0.893\n2)\nLocal CV: 0.902\nPublic LB(TTA): 0.890</p>\n\n<p>Maybe +-2%  is expected.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 842370,
          "author_name": "remicogranne",
          "author_url": "",
          "post_date": "05/11/2020 11:08:00",
          "content": "<p>Yup ! There may be a small difference because of the small testing set.\nTo be clear, the generation of the dataset is extremely time consuming and, in this context, getting a very large training dataset may be important if you want to consider a targeted attack (by spliting the data by JPEG quality factors, embedding schemes, noise level, processing operations etc ...\nWe have noted that this can clearly help a lot ... yet this requires more training samples </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 836417,
      "author_name": "kozistr",
      "author_url": "",
      "post_date": "05/07/2020 01:07:06",
      "content": "<p>```\nArch : EfficientNet-b0\nResolution : 512x512x3\nEpoch : maybe 5 ~ 6 (took about 45 mins per epoch on V100 x 1)\nCV : 0.71 (acc) / 0.81 (weighted auc)\nLB : 0.844</p>\n\n<p>w/ whole dataset\nw/o TTA\n```</p>\n\n<p>------- updated</p>\n\n<p>```\nArch : EfficientNet\nCV : 0.83 (weighted auc)\nLB : 0.896</p>\n\n<p>w/ whole dataset\nw/ TTA\n```</p>\n\n<p>In my case, TTA helps a lot.</p>",
      "votes": null,
      "replies": [
        {
          "id": 836710,
          "author_name": "cgothg",
          "author_url": "",
          "post_date": "05/07/2020 07:06:18",
          "content": "<p>Did you use the whole dataset or just a subset? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 836747,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "05/07/2020 07:56:06",
          "content": "<p>Hmm, the gap between those two are significant. \nHow many time takes to complete on epoch in GPU? (I guess you're using the whole set).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 836816,
          "author_name": "kozistr",
          "author_url": "",
          "post_date": "05/07/2020 09:04:05",
          "content": "<p>yeap! this time, i use the whole data set. it takes about 45 mins per epoch (V100 x 1).\nand i just do early stop (usually stopped at 5 ~ 6 epochs),</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 852225,
          "author_name": "yxtx2018",
          "author_url": "",
          "post_date": "05/18/2020 08:48:35",
          "content": "<p>Hi  <a href=\"/kozistr\">@kozistr</a>, what's the transformations in your TTA？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 852381,
          "author_name": "kozistr",
          "author_url": "",
          "post_date": "05/18/2020 11:46:26",
          "content": "<p>Vertical &amp; Horizontal flips are used!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 836785,
      "author_name": "wuliaokaola",
      "author_url": "",
      "post_date": "05/07/2020 08:28:03",
      "content": "<p>Model : EfficientNet\nLocal CV: 0.853  whole dataset (0.8:0.2)\nPublic LB: 0.889 \nPublic LB(TTA): 0.893</p>",
      "votes": null,
      "replies": [
        {
          "id": 836814,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "05/07/2020 09:02:41",
          "content": "<p>that's great. are you training on GPU or TPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 838298,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "05/08/2020 13:04:23",
          "content": "<p>Training on TPU, because the dataset is a little bit large. </p>\n\n<p>Update:\nLocal CV: 0.877\nPublic LB(TTA): 0.893</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844675,
          "author_name": "cruigo93",
          "author_url": "",
          "post_date": "05/12/2020 19:05:37",
          "content": "<p>Thank you for sharing! Can I ask you a question, did you do it as binary classification or multi? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 844898,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "05/12/2020 23:10:32",
          "content": "<p>Check <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/150359\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 846720,
      "author_name": "vaillant",
      "author_url": "",
      "post_date": "05/14/2020 00:30:46",
      "content": "<p>MixNet\nCV: 0.889 (90/10 split)\nLB: 0.901 \nLB (TTA): 0.894</p>\n\n<p>Not sure why TTA worsens score. I'm doing hflip, vflip, and transpose height/width. </p>",
      "votes": null,
      "replies": [
        {
          "id": 852195,
          "author_name": "yxtx2018",
          "author_url": "",
          "post_date": "05/18/2020 08:22:07",
          "content": "<p>me too, have you figured out the reason?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 857695,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "05/22/2020 20:40:12",
          "content": "<p>it makes sense to look on validation only in this competition) it is ok you get +-1% on LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 866749,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "05/29/2020 16:38:37",
          "content": "<p>Are you using DCT images? If you augment them, I think you might need to correct each 8x8 block so that it is orientated correctly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 867110,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "05/30/2020 02:33:19",
          "content": "<p>I tried it, but it's too slow.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922177,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/09/2020 21:16:16",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> \nensembling two models with TTA gave worse results, but nice boost separately. Didn't figure out why 😯 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 856951,
      "author_name": "chlxyd",
      "author_url": "",
      "post_date": "05/22/2020 07:09:51",
      "content": "<p>efficientnet-b2 80/20 split\ncv(no tta): 913\nlb(with horizon/vertical flip TTA): 920</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 883756,
      "author_name": "etudiant233",
      "author_url": "",
      "post_date": "06/12/2020 22:12:32",
      "content": "<p>So how much time does it take the model to train on the full data set for one epoch? For me it took about 2 hours on the Kaggle GPU but I want to know if that is a reasonable number.</p>",
      "votes": null,
      "replies": [
        {
          "id": 883788,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "06/12/2020 23:37:30",
          "content": "<p>depends on your model architecture. For me it is about one hour on full dataset for eb0. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 901705,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "06/25/2020 16:33:41",
      "content": "<p>efficientnet-b3 80/20 split (hold out)\ncv (w/o tta): 0.926\nlb (w/o tta): 0.925\nlb (with tta): 0.931</p>\n\n<p>How can we reach 0.940+...?</p>",
      "votes": null,
      "replies": [
        {
          "id": 901714,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "06/25/2020 16:37:26",
          "content": "<p>i think blending is the answer:) What is your cv with tta? I have cv with single model w/o tta 0.93, with tta 0.9335, but still lb was 0.93:(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 901749,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "06/25/2020 17:08:22",
          "content": "<p>blending... =(</p>\n\n<p>I tested cv with tta;\ncv (with tta): 0.930</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 902053,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "06/25/2020 22:23:28",
          "content": "<p>My ef-b3 also has a good val score but not high score on LB, do you noticed similar thing? Also by cv do you use 5fold model ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 902073,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "06/25/2020 23:00:00",
          "content": "<p>Nope, most representetive split of fixed images</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 910273,
          "author_name": "fenglrving",
          "author_url": "",
          "post_date": "07/01/2020 04:47:00",
          "content": "<p>What strategies did TTA use? Instead, my LB(with TTA) dropped points</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 910661,
          "author_name": "pranshu29",
          "author_url": "",
          "post_date": "07/01/2020 09:31:19",
          "content": "<p>great to see your results <a href=\"/ren4yu\">@ren4yu</a>  how are training on whole dataset is it on kaggle or on other platform thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922178,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/09/2020 21:18:35",
          "content": "<p><a href=\"/ren4yu\">@ren4yu</a> \nhow many epochs have you trained for E3? I've traind E5 pretty long but didn't get any decent score like you. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922250,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "07/10/2020 00:42:15",
          "content": "<p>I think 150 epochs would be enough. Yes, it requires really long time. I trained on my local GPU machine early in this competition when I was very busy, and I left my machine training a same model for a long time...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924581,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "07/11/2020 14:11:39",
          "content": "<p>🐱 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 915965,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "07/05/2020 08:45:08",
      "content": "<p>Just started:</p>\n\n<p><code>\nefficientnet-b0\nCV: 0.916\nLB: 0.914\n</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 916041,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/05/2020 09:50:46",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Good to see you 😃 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917385,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/06/2020 13:10:16",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> \nAre you training on GPU or TPU? May I know how much time it took per epoch for e0?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917452,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "07/06/2020 14:23:29",
          "content": "<p>GPU locally - it will always depend on your HW</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 916846,
      "author_name": "cooolz",
      "author_url": "",
      "post_date": "07/06/2020 04:43:17",
      "content": "<p>efficientnet-b2</p>\n\n<p>cv 0.923\nlb 0.926 with tta,</p>\n\n<p>And does the big model better?</p>",
      "votes": null,
      "replies": [
        {
          "id": 917388,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/06/2020 13:15:05",
          "content": "<p>have you calculted w/o tta? did it help? \nBig model is different, but training e5/6/7 took forever kind of :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917396,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "07/06/2020 13:23:13",
          "content": "<p>A single b1 model can achieve LB 0.931 w/ TTA.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917470,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "07/06/2020 14:36:30",
          "content": "<p><a href=\"/ren4yu\">@ren4yu</a> Good to know - but I assume there are some tricks necessary for that :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 918138,
          "author_name": "cooolz",
          "author_url": "",
          "post_date": "07/07/2020 02:51:33",
          "content": "<p>i haven't test without tta. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 918139,
          "author_name": "cooolz",
          "author_url": "",
          "post_date": "07/07/2020 02:52:55",
          "content": "<p>Thanks for that. <a href=\"/ren4yu\">@ren4yu</a> \nAnd one more question, how many epochs do your model trained? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922245,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "07/10/2020 00:38:36",
          "content": "<p>I think 150 epochs would be enough. Yes, it requires really long time...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922321,
          "author_name": "cooolz",
          "author_url": "",
          "post_date": "07/10/2020 02:57:41",
          "content": "<p>LOL,  it is really a lone time. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 924558,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/11/2020 13:58:05",
          "content": "<p>150 epoch is too long. \n<a href=\"/cooolz\">@cooolz</a> how many epochs have you trained for this?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 918006,
      "author_name": "usmanr149",
      "author_url": "",
      "post_date": "07/06/2020 21:56:08",
      "content": "<p>Model: SRNet\nEpochs: ~25\nLB: 0.845</p>",
      "votes": null,
      "replies": [
        {
          "id": 921182,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "07/09/2020 05:41:01",
          "content": "<p>You implement your own SRNet? I also tried SRNet but that only allows me fed 2 images one batch (Kaggle notebook). May I ask what's your batch_size?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 920981,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "07/09/2020 01:59:50",
      "content": "<p>efficientnetB2\nsingle fold\nCV: 0.9079 LB: 0.907 (no tta)</p>",
      "votes": null,
      "replies": [
        {
          "id": 920992,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/09/2020 02:08:57",
          "content": "<p>TTA improvement is about +0.004 (for horizonal and veritical flip)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 922180,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/09/2020 21:22:27",
          "content": "<p>It's somewhat misleading. It's not about your model score, but with the same model some other reporting better score. So, I think it's not proper to say only the base model, like others it has to do something with the head of this base model.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 926132,
      "author_name": "authman",
      "author_url": "",
      "post_date": "07/12/2020 14:17:07",
      "content": "<p>Had originally experimented with effnet b0 based on forums recommendations. This was an excellent idea. It's super fast and reasonable. Started getting results that looked okay (they matched the results of the best public kernel, even though b0) so I got thirsty/greedy and made the huge mistake of starting to train a b3 model. Each epoch takes about an hour and some change on my dual 2080Ti machine. Since time is ticking, decided to make a sub off of current best epoch (31 out of 34 completed). From my observations, it is not overfit yet and can still maybe train an additional 15-20 epochs(?) to secure maybe an additional +0.005 I think. This is a single fold pytorch model trained on 80% of the data (GKF=5). I also have 3 other architectures I've designed but hadn't had compute resources to test up yet.</p>\n\n<p>B3 31 epoch\nTrain: CV 0.91857\nVal: CV 0.92084\nVal w/ TTA: 0.9289\nSub w/ TTA: 0.929</p>\n\n<p>B0 37 epoch (same architectural changes + but w/ less train augmentation and no TTA):\nTrain: CV 0.91948\nVal: CV 0.92011\nSub w/o TTA: 0.917</p>",
      "votes": null,
      "replies": [
        {
          "id": 926213,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "07/12/2020 15:21:59",
          "content": "<p>And how does B0 with TTA on CV/LB?</p>\n\n<p>CV 0.919 without TTA on B0 is pretty good, seems to match your B3 (even if trained a bit longer).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926244,
          "author_name": "authman",
          "author_url": "",
          "post_date": "07/12/2020 15:40:11",
          "content": "<blockquote>\n  <p>And how does B0 with TTA on CV/LB?</p>\n</blockquote>\n\n<p>No clue. I still have the checkpoints + notebook from the b0 model so I guess I could use kaggle kernel to run TTA on it(?).</p>\n\n<p>B3 model is still training locally + val CV just hit 0.92252 at epoch 35. Submitted and score is 0.931-LB. I think it might be time to figure out how this whole tpu thing works. My understanding is that the TPU node is hosted physically close to the computer doing training, ideally on the same subnet. The major bottleneck I'm seeing on my machine is decoding of JPEG. I tried caching the decoded jpeg's as .npy and as .pkl but then my SSD's light would come on and I'd hit transfer speed bottlenecks instead of decoding speed bottlenecks.</p>\n\n<p>Are you using TPU? Are you aware of the throughput of google's SSD? If they aren't remarkable, then I'd assume that the same transfer bottlenecks would occur, limiting the amount of data that could be pumped into the TPU........</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932105,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "07/16/2020 17:43:42",
          "content": "<p>If you use pinned memory and multiple workers in your dataloader, JPG decoding shouldn't be a problem. Are you working on Linux?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932160,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "07/16/2020 19:05:35",
          "content": "<p><a href=\"/authman\">@authman</a> TTA boosting 8 pts? Thats quite a bit...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932192,
          "author_name": "authman",
          "author_url": "",
          "post_date": "07/16/2020 19:49:50",
          "content": "<p><a href=\"/bigironsphere\">@bigironsphere</a> Was not using pinned memory. I'll give it a try, TY.</p>\n\n<p><a href=\"/philippsinger\">@philippsinger</a> ikr? TTA seems to work particularly well with EffNet. I've experimented with both B0 and B3. However, the exact same TTA with some of the other networks I've seen people use from the External Data Thread like PyConvNet, RexNet, ResNet, etc. etc. don't have as spectacular results. For example, yesterday's run:</p>\n\n<p>Epoch 42\nTrain: 0.92658\nVal: 0.92563\nVal TTA: 0.92819\nLB: 0.927</p>\n\n<p>I haven't started looking at model correlation yet. Will do so at the last moment when there's no more time to train.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932196,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "07/16/2020 19:50:54",
          "content": "<p>0.930 -&gt; 0.926 (with TTA) for me, EfficientNet-B1</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932199,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/16/2020 19:57:30",
          "content": "<p><a href=\"/authman\">@authman</a> \nWhat's the result of your experiment on PyConvNet? Is it from this <a href=\"https://github.com/iduta/pyconv\">implementation</a>? Have you tried on GhostNet or Xception?</p>\n\n<p><a href=\"/vaillant\">@vaillant</a> \nTTA works well in my cases, just horizontal/vertical flip average. May I know what you've followed for this? I'm worried about TTA, will it overfit on private set!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932202,
          "author_name": "authman",
          "author_url": "",
          "post_date": "07/16/2020 20:03:40",
          "content": "<p>Fascinating. Perhaps the better base model you have, the less TTA works. We'll need some of the +top20's to weigh in.</p>\n\n<p>Yup, that was the PCN implementation I used.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932868,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "07/17/2020 11:07:01",
          "content": "<blockquote>\n  <p><strong>Ian Pan wrote:</strong></p>\n  \n  <p>0.930 -&gt; 0.926 (with TTA) for me, EfficientNet-B1</p>\n</blockquote>\n\n<p>our observation is that the relationship between CV/LB is fold-dependent.\nIn our 5 folds, we have some 3 folds get better CV than LB, the other two the other way around. and that is quite consistent both to different models, and among different iterations of the same model</p>\n\n<p>It might be a bit of a lottery in the end, because it is quite likely the private LB would yield relatively different score as compared to public LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932882,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/17/2020 11:14:19",
          "content": "<p><a href=\"/yifanxie\">@yifanxie</a> \nI have also observed similar stuff on fold-dependencies stuff. The ensemble of different models on the same fold with TTA improves significant but little on the different folds. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932891,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/17/2020 11:17:46",
          "content": "<p>However, In our observation, small variants of model score much better than their larger variants with a reasonable config for both. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932902,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "07/17/2020 11:26:10",
          "content": "<blockquote>\n  <p><strong>M.Innat wrote:</strong></p>\n  \n  <p>However, In our observation, small variants of model score much better than their larger variants with a reasonable config for both. </p>\n</blockquote>\n\n<p>If I understand correctly, you are saying it is easier to get good score with smaller models (i.e. B0/B1) than larger models (i.e. B6/B7)?</p>\n\n<p>Our larger models are still in the oven, so we will see - our latest model is due to finished 2 hours before the deadline assuming no power outage 😌 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932927,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/17/2020 11:43:27",
          "content": "<p><a href=\"/yifanxie\">@yifanxie</a> \nKinda. Larger variants of EfficientNet (B7) didn't give a better score than its smaller variants. Also surprisingly lower than MixNet. As <a href=\"/ren4yu\">@ren4yu</a> <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/148162#917396\">mentioned earlier</a>, is kind of holds true here. But, it does not give any conclusion, if I'm not wrong, efficientnet is somewhat resolution-dependent. So, I was wondering if E7 needs a resolution &gt; than 512.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932929,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "07/17/2020 11:43:51",
          "content": "<blockquote>\n  <p>our latest model is due to finished 2 hours before the deadline assuming no power outage</p>\n</blockquote>\n\n<p>Living on the edges 😎 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932955,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "07/17/2020 11:56:21",
          "content": "<blockquote>\n  <p><strong>Ahmet Erdem wrote:</strong>\n  Living on the edges 😎 </p>\n</blockquote>\n\n<p>more like beggar can't be chooser 🙌 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932975,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "07/17/2020 12:02:55",
          "content": "<p>Is this competition more computationally expensive than DeepFake one btw?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932992,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/17/2020 12:09:36",
          "content": "<p>From me, I wouldn't have any chances to compete with this <strong>computation hunger</strong> competition without my company support. Feel lucky to be a part of such corporations. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932998,
          "author_name": "yifanxie",
          "author_url": "",
          "post_date": "07/17/2020 12:14:17",
          "content": "<blockquote>\n  <p><strong>Ahmet Erdem wrote:</strong>\n  Is this competition more computationally expensive than DeepFake one btw?</p>\n</blockquote>\n\n<p>I am personally using 3~4x more GPU resources 😿 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933055,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "07/17/2020 13:01:52",
          "content": "<p>Hopefully shakeup is kind to us. 😂 </p>\n\n<p>I was referring to my score on public LB with/without TTA. I'm using anokas' weighted AUC metric, and I had to quick fix a bug that would occur. However this bug gives me incorrect AUC values locally after a certain point (goes from 0.91/0.92-&gt;0.35), so not even sure what my CV is...</p>\n\n<p>For vanilla EfficientNet models, I found that larger models do better than smaller models. It will be interesting to see everyone's approaches at competition end.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933069,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/17/2020 13:08:34",
          "content": "<p>I feel really curious to see the solution of the final top teams. And yes, the shakeup is kind to us 😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933510,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "07/17/2020 18:29:58",
          "content": "<p>Around 1 month ago, I got sth like: improve from local 9245 to 9294, but got LB 932 and 927 respectively. 😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933515,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/17/2020 18:34:49",
          "content": "<p>That's long...😄 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 937375,
      "author_name": "tivfrvqhs5",
      "author_url": "",
      "post_date": "07/21/2020 01:18:49",
      "content": "<p>B7 w/bitmix + TTA+ train on all data: 0.937 public/0.925 private</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937425,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "07/21/2020 02:33:57",
      "content": "<p>B7 + TTA + train on all data + finetune upper layers = 0.945 public / 0.926 private.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937446,
      "author_name": "wuliaokaola",
      "author_url": "",
      "post_date": "07/21/2020 03:04:37",
      "content": "<p>B6 + TTA \nLocal / Public / Private: 0.940 / 0.940 / 0.929</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 938020,
      "author_name": "nuller",
      "author_url": "",
      "post_date": "07/21/2020 09:54:40",
      "content": "<p>effb4, single fold +  TTA (rot90, flips)\nlocal/public/private: 0.931/0.930/0.926</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "831429": "Let's share some insight. =)\n\n**Baseline** (w/o full data set)\n```\nModel: EfficientNet\nEpoch: 5\nCV: 0.71\nLB: 0.75\n```\n\n## Update\n\nnearly full data set.\n```\nModel: EfficientNet\nEpoch: ~15\nCV: 0.88\nLB: 0.84\n```\n\n## Update\n\n```\nModel: MixNet\nCV: 0.913\nLB: 0.910\n```",
    "832371": "```\nModel : EfficientNet-b3 (w/ some augs)\nResolution : 512x512x3 (rgb)\nEpoch : 12 (w/o EarlyStop)\nCV : 0.62 (90/10 split)\nLB : 0.801\n\nw/o full dataset\nw/o TTA\n```\nLB fluctuates depending on the seed. (it's not good)",
    "832453": "Are you training on TPU or GPU?",
    "832485": "training on TPU! \ncurrently, I'm trying to use higher resolution &amp; batch size, GPU can't handle it cuz of OOM :( so I use TPU.",
    "832491": "I am also facing the seed issue. However, I am also working on TPU but I think, the usage time limitation is a big barrier, on top that some unexpected issue related to TPU causes time kills :(",
    "832851": "I 'm also a.l.s.o. facing the seed issue. :(\n\nI already used TPU quota in this week.",
    "834033": "efficientnet-b0, 512, 512, 3\nlocal val: 0.890, public LB: 0.895",
    "834039": "Your CV and LB correlated. Have you tried on the whole dataset?",
    "834189": "It's worth noting that the public LB is really small, only 1000 samples, so some fluctuation +-1% at least t is expected.\n\nI wonder why the organisers didn't make the test set bigger given they were able to obtain 75,000 images for training (different sources?)",
    "834190": "Wow, that's a lot better than I got with efficientnet (0.853 CV, 0.880 LB)",
    "834316": "Yes, actually I split the validation set first, and gradually increase the number of training samples. During this process, in most of the time cv and lb are correlated.",
    "834320": "I also got larger gap during my old experiments, FYI:\nlocal 0.845, LB 0.869\nlocal 0.860, LB 0.884",
    "836417": "```\nArch : EfficientNet-b0\nResolution : 512x512x3\nEpoch : maybe 5 ~ 6 (took about 45 mins per epoch on V100 x 1)\nCV : 0.71 (acc) / 0.81 (weighted auc)\nLB : 0.844\n\nw/ whole dataset\nw/o TTA\n```\n\n------- updated\n\n```\nArch : EfficientNet\nCV : 0.83 (weighted auc)\nLB : 0.896\n\nw/ whole dataset\nw/ TTA\n```\n\nIn my case, TTA helps a lot.",
    "836710": "Did you use the whole dataset or just a subset?",
    "836747": "Hmm, the gap between those two are significant. \nHow many time takes to complete on epoch in GPU? (I guess you're using the whole set).",
    "836785": "Model : EfficientNet\nLocal CV: 0.853  whole dataset (0.8:0.2)\nPublic LB: 0.889 \nPublic LB(TTA): 0.893",
    "836814": "that's great. are you training on GPU or TPU?",
    "836816": "yeap! this time, i use the whole data set. it takes about 45 mins per epoch (V100 x 1).\nand i just do early stop (usually stopped at 5 ~ 6 epochs),",
    "838298": "Training on TPU, because the dataset is a little bit large. \n\nUpdate:\nLocal CV: 0.877\nPublic LB(TTA): 0.893",
    "841804": "I have two models with the same split.\n1)\nLocal CV: 0.877\nPublic LB(TTA): 0.893\n2)\nLocal CV: 0.902\nPublic LB(TTA): 0.890\n\nMaybe +-2%  is expected.",
    "842363": "Well done !\nI strongly invite you to test the following  (see my post [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/148919) and [there](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/147494) ):\n- Try multiclass for idenfying each embedding scheme individually\n- Try split the dataset in terms of JPEG quality factor (there is only 3 of them, you can easily guess by look at quantization table)\n- Increase the training set (possibly using data augmentation , [see also this paper that try many possible ways and show there efficiency](https://hal-lirmm.ccsd.cnrs.fr/lirmm-02559838/file/IHMMSec-2016_Yedroudj_Chaumont_Comby_Amara_Bas_Pixels-off.pdf))\n- Try incorporating DCT coefficients\n- and, of course, other tricks that you guys know much better than me such as using several models and merge them (as already proposed by merging efficient net b3 and b7 ...)",
    "842370": "Yup ! There may be a small difference because of the small testing set.\nTo be clear, the generation of the dataset is extremely time consuming and, in this context, getting a very large training dataset may be important if you want to consider a targeted attack (by spliting the data by JPEG quality factors, embedding schemes, noise level, processing operations etc ...\nWe have noted that this can clearly help a lot ... yet this requires more training samples",
    "842377": "Here's a snippet to split by JPEG quality factor (it is a public data, it is simply in the header of the JPEG files ....)\n`for file in train_files[start: end]:`\n`____fullCover = folders[3] + '/' + file`\n`____randFolder = np.random.randint(0,3)`\n`____full = folders[randFolder] + '/' + file`\n`____imgJPEG = jio.read(path+folders[0]+'/'+file)`\n`____if( imgJPEG.quant_tables[0][0,0] == 2):`\n`________cover_paths95.append(fullCover)`\n`________steg_paths95.append(full)`\n`____elif( imgJPEG.quant_tables[0][0,0] == 3):`\n`________cover_paths90.append(fullCover)`\n`________steg_paths90.append(full)`\n`____elif( imgJPEG.quant_tables[0][0,0] == 8):`\n`________cover_paths75.append(fullCover)`\n`________steg_paths75.append(full)`",
    "842386": "remicogranne thanks, it helps indeed. :)",
    "842565": "Thanks!",
    "844675": "Thank you for sharing! Can I ask you a question, did you do it as binary classification or multi?",
    "844898": "Check [here](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/150359)",
    "846720": "MixNet\nCV: 0.889 (90/10 split)\nLB: 0.901 \nLB (TTA): 0.894\n\nNot sure why TTA worsens score. I'm doing hflip, vflip, and transpose height/width.",
    "852195": "me too, have you figured out the reason?",
    "852225": "Hi  @kozistr, what's the transformations in your TTA？",
    "852381": "Vertical &amp; Horizontal flips are used!",
    "856951": "efficientnet-b2 80/20 split\ncv(no tta): 913\nlb(with horizon/vertical flip TTA): 920",
    "857695": "it makes sense to look on validation only in this competition) it is ok you get +-1% on LB",
    "866749": "Are you using DCT images? If you augment them, I think you might need to correct each 8x8 block so that it is orientated correctly",
    "867110": "I tried it, but it's too slow.",
    "872588": "Hello! 😁 And what is the \"jio\" library? I can't find it, so I am blindly assuming is a shortcut from an actual name...",
    "873234": "andradaolteanu `jpegio`",
    "883756": "So how much time does it take the model to train on the full data set for one epoch? For me it took about 2 hours on the Kaggle GPU but I want to know if that is a reasonable number.",
    "883788": "depends on your model architecture. For me it is about one hour on full dataset for eb0.",
    "901705": "efficientnet-b3 80/20 split (hold out)\ncv (w/o tta): 0.926\nlb (w/o tta): 0.925\nlb (with tta): 0.931\n\nHow can we reach 0.940+...?",
    "901714": "i think blending is the answer:) What is your cv with tta? I have cv with single model w/o tta 0.93, with tta 0.9335, but still lb was 0.93:(",
    "901749": "blending... =(\n\nI tested cv with tta;\ncv (with tta): 0.930",
    "902053": "My ef-b3 also has a good val score but not high score on LB, do you noticed similar thing? Also by cv do you use 5fold model ?",
    "902073": "Nope, most representetive split of fixed images",
    "906126": "kozistr Hi! Are you still using TPU? If so, are you using full dataset and full resolution? How did you bypass the memory issue.",
    "909500": "Hi @tonychenxyz !\nNope, i'm currently using GPU w/ full dataset, resolution. \nYou can read the image files at runtime instead of loading full dataset on the memory at once.\nIt can solve the memory issue, but I/O read will cost. (but, in this case, it's not a big issue though :) )",
    "910273": "What strategies did TTA use? Instead, my LB(with TTA) dropped points",
    "910661": "great to see your results @ren4yu  how are training on whole dataset is it on kaggle or on other platform thanks",
    "915965": "Just started:\n\n```\nefficientnet-b0\nCV: 0.916\nLB: 0.914\n```",
    "916041": "philippsinger Good to see you 😃",
    "916846": "efficientnet-b2\n\ncv 0.923\nlb 0.926 with tta,\n\nAnd does the big model better?",
    "917385": "philippsinger \nAre you training on GPU or TPU? May I know how much time it took per epoch for e0?",
    "917388": "have you calculted w/o tta? did it help? \nBig model is different, but training e5/6/7 took forever kind of :(",
    "917396": "A single b1 model can achieve LB 0.931 w/ TTA.",
    "917452": "GPU locally - it will always depend on your HW",
    "917470": "ren4yu Good to know - but I assume there are some tricks necessary for that :)",
    "918006": "Model: SRNet\nEpochs: ~25\nLB: 0.845",
    "918138": "i haven't test without tta.",
    "918139": "Thanks for that. @ren4yu \nAnd one more question, how many epochs do your model trained?",
    "920981": "efficientnetB2\nsingle fold\nCV: 0.9079 LB: 0.907 (no tta)",
    "920992": "TTA improvement is about +0.004 (for horizonal and veritical flip)",
    "921182": "You implement your own SRNet? I also tried SRNet but that only allows me fed 2 images one batch (Kaggle notebook). May I ask what's your batch_size?",
    "922177": "vaillant \nensembling two models with TTA gave worse results, but nice boost separately. Didn't figure out why 😯",
    "922178": "ren4yu \nhow many epochs have you trained for E3? I've traind E5 pretty long but didn't get any decent score like you.",
    "922180": "It's somewhat misleading. It's not about your model score, but with the same model some other reporting better score. So, I think it's not proper to say only the base model, like others it has to do something with the head of this base model.",
    "922245": "I think 150 epochs would be enough. Yes, it requires really long time...",
    "922250": "I think 150 epochs would be enough. Yes, it requires really long time. I trained on my local GPU machine early in this competition when I was very busy, and I left my machine training a same model for a long time...",
    "922321": "LOL,  it is really a lone time.",
    "924558": "150 epoch is too long. \n@cooolz how many epochs have you trained for this?",
    "924581": "🐱",
    "926132": "Had originally experimented with effnet b0 based on forums recommendations. This was an excellent idea. It's super fast and reasonable. Started getting results that looked okay (they matched the results of the best public kernel, even though b0) so I got thirsty/greedy and made the huge mistake of starting to train a b3 model. Each epoch takes about an hour and some change on my dual 2080Ti machine. Since time is ticking, decided to make a sub off of current best epoch (31 out of 34 completed). From my observations, it is not overfit yet and can still maybe train an additional 15-20 epochs(?) to secure maybe an additional +0.005 I think. This is a single fold pytorch model trained on 80% of the data (GKF=5). I also have 3 other architectures I've designed but hadn't had compute resources to test up yet.\n\nB3 31 epoch\nTrain: CV 0.91857\nVal: CV 0.92084\nVal w/ TTA: 0.9289\nSub w/ TTA: 0.929\n\n\nB0 37 epoch (same architectural changes + but w/ less train augmentation and no TTA):\nTrain: CV 0.91948\nVal: CV 0.92011\nSub w/o TTA: 0.917",
    "926213": "And how does B0 with TTA on CV/LB?\n\nCV 0.919 without TTA on B0 is pretty good, seems to match your B3 (even if trained a bit longer).",
    "926244": "&gt; And how does B0 with TTA on CV/LB?\n\nNo clue. I still have the checkpoints + notebook from the b0 model so I guess I could use kaggle kernel to run TTA on it(?).\n\nB3 model is still training locally + val CV just hit 0.92252 at epoch 35. Submitted and score is 0.931-LB. I think it might be time to figure out how this whole tpu thing works. My understanding is that the TPU node is hosted physically close to the computer doing training, ideally on the same subnet. The major bottleneck I'm seeing on my machine is decoding of JPEG. I tried caching the decoded jpeg's as .npy and as .pkl but then my SSD's light would come on and I'd hit transfer speed bottlenecks instead of decoding speed bottlenecks.\n\nAre you using TPU? Are you aware of the throughput of google's SSD? If they aren't remarkable, then I'd assume that the same transfer bottlenecks would occur, limiting the amount of data that could be pumped into the TPU........",
    "932105": "If you use pinned memory and multiple workers in your dataloader, JPG decoding shouldn't be a problem. Are you working on Linux?",
    "932160": "authman TTA boosting 8 pts? Thats quite a bit...",
    "932192": "bigironsphere Was not using pinned memory. I'll give it a try, TY.\n\n@philippsinger ikr? TTA seems to work particularly well with EffNet. I've experimented with both B0 and B3. However, the exact same TTA with some of the other networks I've seen people use from the External Data Thread like PyConvNet, RexNet, ResNet, etc. etc. don't have as spectacular results. For example, yesterday's run:\n\nEpoch 42\nTrain: 0.92658\nVal: 0.92563\nVal TTA: 0.92819\nLB: 0.927\n\nI haven't started looking at model correlation yet. Will do so at the last moment when there's no more time to train.",
    "932196": "0.930 -&gt; 0.926 (with TTA) for me, EfficientNet-B1",
    "932199": "authman \nWhat's the result of your experiment on PyConvNet? Is it from this [implementation](https://github.com/iduta/pyconv)? Have you tried on GhostNet or Xception?\n\n@vaillant \nTTA works well in my cases, just horizontal/vertical flip average. May I know what you've followed for this? I'm worried about TTA, will it overfit on private set!",
    "932202": "Fascinating. Perhaps the better base model you have, the less TTA works. We'll need some of the +top20's to weigh in.\n\nYup, that was the PCN implementation I used.",
    "932868": "&gt; **Ian Pan wrote:**\n&gt; \n&gt; 0.930 -&gt; 0.926 (with TTA) for me, EfficientNet-B1\n\nour observation is that the relationship between CV/LB is fold-dependent.\nIn our 5 folds, we have some 3 folds get better CV than LB, the other two the other way around. and that is quite consistent both to different models, and among different iterations of the same model\n\nIt might be a bit of a lottery in the end, because it is quite likely the private LB would yield relatively different score as compared to public LB",
    "932882": "yifanxie \nI have also observed similar stuff on fold-dependencies stuff. The ensemble of different models on the same fold with TTA improves significant but little on the different folds.",
    "932891": "However, In our observation, small variants of model score much better than their larger variants with a reasonable config for both.",
    "932902": "&gt; **M.Innat wrote:**\n&gt; \n&gt; However, In our observation, small variants of model score much better than their larger variants with a reasonable config for both. \n\nIf I understand correctly, you are saying it is easier to get good score with smaller models (i.e. B0/B1) than larger models (i.e. B6/B7)?\n\nOur larger models are still in the oven, so we will see - our latest model is due to finished 2 hours before the deadline assuming no power outage 😌",
    "932927": "yifanxie \nKinda. Larger variants of EfficientNet (B7) didn't give a better score than its smaller variants. Also surprisingly lower than MixNet. As @ren4yu [mentioned earlier](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/148162#917396), is kind of holds true here. But, it does not give any conclusion, if I'm not wrong, efficientnet is somewhat resolution-dependent. So, I was wondering if E7 needs a resolution &gt; than 512.",
    "932929": "&gt; our latest model is due to finished 2 hours before the deadline assuming no power outage\n\nLiving on the edges 😎",
    "932955": "&gt; **Ahmet Erdem wrote:**\n&gt; Living on the edges 😎 \n\nmore like beggar can't be chooser 🙌",
    "932975": "Is this competition more computationally expensive than DeepFake one btw?",
    "932992": "From me, I wouldn't have any chances to compete with this **computation hunger** competition without my company support. Feel lucky to be a part of such corporations.",
    "932998": "&gt; **Ahmet Erdem wrote:**\n&gt; Is this competition more computationally expensive than DeepFake one btw?\n\nI am personally using 3~4x more GPU resources 😿",
    "933055": "Hopefully shakeup is kind to us. 😂 \n\nI was referring to my score on public LB with/without TTA. I'm using anokas' weighted AUC metric, and I had to quick fix a bug that would occur. However this bug gives me incorrect AUC values locally after a certain point (goes from 0.91/0.92-&gt;0.35), so not even sure what my CV is...\n\nFor vanilla EfficientNet models, I found that larger models do better than smaller models. It will be interesting to see everyone's approaches at competition end.",
    "933069": "I feel really curious to see the solution of the final top teams. And yes, the shakeup is kind to us 😄",
    "933510": "Around 1 month ago, I got sth like: improve from local 9245 to 9294, but got LB 932 and 927 respectively. 😂",
    "933515": "That's long...😄",
    "937375": "B7 w/bitmix + TTA+ train on all data: 0.937 public/0.925 private",
    "937425": "B7 + TTA + train on all data + finetune upper layers = 0.945 public / 0.926 private.",
    "937446": "B6 + TTA \nLocal / Public / Private: 0.940 / 0.940 / 0.929",
    "937970": "B4 Fold4 (public GKF notebook) local .931 private .925 (enough for 24th place)\nB4 (5 folds) local .930 private .924 (enough for 30th place)\n\nFinally, I enjoyed learning about RegNetY, it was fun to train, and does so quite \"quickly\" (28epo@34min, single GPU)\nRegNetY-8 (5 folds) +TTA local .924 private .920 (enough for 70th place)",
    "938020": "effb4, single fold +  TTA (rot90, flips)\nlocal/public/private: 0.931/0.930/0.926"
  },
  "source": "meta"
}