{
  "id": 123198,
  "title": "Best single model",
  "url": "/competitions/bengaliai-cv19/discussion/123198",
  "author_name": "DrHB",
  "post_date": "2019-12-25T16:14:24.096000",
  "votes": 148,
  "comment_count": 408,
  "views": 0,
  "content": "<p>Just starting a common competition thread =) \nWhat is your current best single model?\nMy:</p>\n\n<p>```\nmodel: xresnet18(pretrained=False)\nimg: 1 channel \nimg_sz: 128\nsplit: random (80/20)\noptim: Adam\nepoch: 15\nsched: Cosine decay</p>\n\n<p>CV : 0.9630\nLB:  0.9611\n```</p>",
  "messages": [
    {
      "id": 703111,
      "postDate": "2019-12-25T16:14:24.097Z",
      "content": "<p>Just starting a common competition thread =) \nWhat is your current best single model?\nMy:</p>\n\n<p>```\nmodel: xresnet18(pretrained=False)\nimg: 1 channel \nimg_sz: 128\nsplit: random (80/20)\noptim: Adam\nepoch: 15\nsched: Cosine decay</p>\n\n<p>CV : 0.9630\nLB:  0.9611\n```</p>",
      "rawMarkdown": "Just starting a common competition thread =) \nWhat is your current best single model?\nMy:\n\n```\nmodel: xresnet18(pretrained=False)\nimg: 1 channel \nimg_sz: 128\nsplit: random (80/20)\noptim: Adam\nepoch: 15\nsched: Cosine decay\n\nCV : 0.9630\nLB:  0.9611\n```",
      "votes": 145
    },
    {
      "id": 718926,
      "postDate": "2020-01-15T01:03:22.753Z",
      "content": "<p>[Update]\n<code>\nmodel: se-resnext50 with mixup + cutmix\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.988011\nLB:  0.9790\n</code>\n<code>python\nif np.random.rand()&lt;0.5: do mixup\nelse: do cutmix\n</code></p>",
      "rawMarkdown": "[Update]\n```\nmodel: se-resnext50 with mixup + cutmix\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.988011\nLB:  0.9790\n```\n```python\nif np.random.rand()&lt;0.5: do mixup\nelse: do cutmix\n```",
      "votes": 37,
      "replies": [
        {
          "id": 719043,
          "postDate": "2020-01-15T04:04:42.433Z",
          "content": "<p>Thanks for the update. It is really helpful.</p>\n\n<p>How do you process input image? Is gray-scaled 3x137x236 image normalized? </p>",
          "rawMarkdown": "Thanks for the update. It is really helpful.\n\nHow do you process input image? Is gray-scaled 3x137x236 image normalized? \n\n"
        },
        {
          "id": 719061,
          "postDate": "2020-01-15T04:43:04.863Z",
          "content": "<p>With my training loss being 0.01 and cross validation loss as 0.10, there is some overfitting. Generalisation will improve results</p>",
          "rawMarkdown": "With my training loss being 0.01 and cross validation loss as 0.10, there is some overfitting. Generalisation will improve results"
        },
        {
          "id": 719086,
          "postDate": "2020-01-15T05:49:26.897Z",
          "content": "<p><a href=\"/bibek777\">@bibek777</a>  I think you have saved an average of a week or two of time for many people ;) by your wonderful posts</p>",
          "rawMarkdown": "@bibek777  I think you have saved an average of a week or two of time for many people ;) by your wonderful posts",
          "votes": 1
        },
        {
          "id": 719093,
          "postDate": "2020-01-15T05:57:11.130Z",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> I get your concern. I will be careful from next time. </p>",
          "rawMarkdown": "@haqishen I get your concern. I will be careful from next time. "
        },
        {
          "id": 719100,
          "postDate": "2020-01-15T06:09:07.707Z",
          "content": "<p>Sharing ideas that can reach gold zone is always controversial (but it's ok now since we have 2 month left), anyway, congratulation for you have actually reached 0.98 zone (you only need to modify image size to 224x224 for truly reaching that line)</p>",
          "rawMarkdown": "Sharing ideas that can reach gold zone is always controversial (but it's ok now since we have 2 month left), anyway, congratulation for you have actually reached 0.98 zone (you only need to modify image size to 224x224 for truly reaching that line)",
          "votes": 9
        },
        {
          "id": 719102,
          "postDate": "2020-01-15T06:13:04.893Z",
          "content": "<p>Now I know what I will do tonite🙏 🙏 🙏 </p>",
          "rawMarkdown": "Now I know what I will do tonite🙏 🙏 🙏 ",
          "votes": 1
        },
        {
          "id": 719133,
          "postDate": "2020-01-15T07:02:22.797Z",
          "content": "<p><a href=\"/bibek777\">@bibek777</a> thanks for sharing. How are you able to test things out so quickly? How long is each epoch taking for you?</p>",
          "rawMarkdown": "@bibek777 thanks for sharing. How are you able to test things out so quickly? How long is each epoch taking for you?",
          "votes": 2
        },
        {
          "id": 719232,
          "postDate": "2020-01-15T09:10:38.823Z",
          "content": "<blockquote>\n  <p>How are you able to test things out so quickly? </p>\n</blockquote>\n\n<p>when you are in top10, you spend your time kaggling rather than timetravelling😜 😜 😜 </p>",
          "rawMarkdown": "&gt;  How are you able to test things out so quickly? \n\nwhen you are in top10, you spend your time kaggling rather than timetravelling😜 😜 😜 ",
          "votes": 2
        },
        {
          "id": 719467,
          "postDate": "2020-01-15T14:25:54.913Z",
          "content": "<p><a href=\"/timetraveller98\">@timetraveller98</a>  , I am in his team , and its hard for me to keep up ..lol :P </p>",
          "rawMarkdown": "@timetraveller98  , I am in his team , and its hard for me to keep up ..lol :P ",
          "votes": 2
        },
        {
          "id": 719662,
          "postDate": "2020-01-15T17:54:08.770Z",
          "content": "<p>Thx for sharing. Did yоu use Sampler?</p>",
          "rawMarkdown": "Thx for sharing. Did yоu use Sampler?",
          "votes": 1
        },
        {
          "id": 732794,
          "postDate": "2020-01-30T09:21:58.763Z",
          "content": "<p>Thanks.</p>",
          "rawMarkdown": "Thanks."
        }
      ]
    },
    {
      "id": 717760,
      "postDate": "2020-01-13T15:34:28.717Z",
      "content": "<p>The submission of\n<code>LB 0.9875</code>\nwas a single fold model which scored on my local experiment by\n<code>CV 0.996792</code></p>\n\n<p>Hope this will be motivation for you guys to improve the performance of single model, good luck!</p>",
      "rawMarkdown": "The submission of\n`LB 0.9875`\nwas a single fold model which scored on my local experiment by\n`CV 0.996792`\n\n\nHope this will be motivation for you guys to improve the performance of single model, good luck!",
      "votes": 36,
      "replies": [
        {
          "id": 717781,
          "postDate": "2020-01-13T16:02:43.820Z",
          "content": "<p>Great. I'm guessing your current LB score is by ensembling! What was your baseline setup, like model architecture, optimizer or validation strategies? What about the data augmentation approach, any advantages with that? A single model is scoring really well, and I doubt whether we need to the ensemble! 🙄 </p>",
          "rawMarkdown": "Great. I'm guessing your current LB score is by ensembling! What was your baseline setup, like model architecture, optimizer or validation strategies? What about the data augmentation approach, any advantages with that? A single model is scoring really well, and I doubt whether we need to the ensemble! 🙄 ",
          "votes": 3
        },
        {
          "id": 717786,
          "postDate": "2020-01-13T16:09:21.653Z",
          "content": "<p>Why not read some posts from previous image competitions like these: \n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065</a>\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107926\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107926</a>\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117210\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117210</a>\n<a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118080\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/118080</a>\n<a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543</a></p>\n\n<p>I'm not saying the ideas in those posts are useful for this competition too (you can check by yourself), but saying that you can learn how they were thinking during competition.</p>",
          "rawMarkdown": "Why not read some posts from previous image competitions like these: \nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107926\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117210\nhttps://www.kaggle.com/c/understanding_cloud_organization/discussion/118080\nhttps://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543\n\nI'm not saying the ideas in those posts are useful for this competition too (you can check by yourself), but saying that you can learn how they were thinking during competition.",
          "votes": 21
        },
        {
          "id": 717788,
          "postDate": "2020-01-13T16:13:12.303Z",
          "content": "<p>I went through this one for sure 🖤 😃 \n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107987\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107987</a></p>",
          "rawMarkdown": "I went through this one for sure 🖤 😃 \nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107987",
          "votes": 1
        },
        {
          "id": 717792,
          "postDate": "2020-01-13T16:19:11.490Z",
          "content": "<p>😂 😂 Thank you for bringing it up.\nIt was a hard competition to me.</p>",
          "rawMarkdown": "😂 😂 Thank you for bringing it up.\nIt was a hard competition to me.",
          "votes": 1
        },
        {
          "id": 717793,
          "postDate": "2020-01-13T16:20:21.107Z",
          "content": "<p>ahhh APTOS, good times.... was so stressful =) </p>",
          "rawMarkdown": "ahhh APTOS, good times.... was so stressful =) ",
          "votes": 4
        },
        {
          "id": 717797,
          "postDate": "2020-01-13T16:24:08.240Z",
          "content": "<p>Yea... Exactly it was\nSo let me bring up your excellent post as well 😄 \n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108030\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108030</a></p>",
          "rawMarkdown": "Yea... Exactly it was\nSo let me bring up your excellent post as well 😄 \nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108030",
          "votes": 2
        },
        {
          "id": 718004,
          "postDate": "2020-01-13T21:42:15.690Z",
          "content": "<p>Did you balance your training data?</p>",
          "rawMarkdown": "Did you balance your training data?"
        }
      ]
    },
    {
      "id": 723165,
      "postDate": "2020-01-19T15:19:55.223Z",
      "content": "<p>CV: 0.9913\nLB: 0.9848</p>\n\n<p>though I still haven't reached a single-model score as high as <a href=\"/haqishen\">@haqishen</a> but I haven't used any TTA and big boys like seresnext or high efficientnet/densenet due to free time and computational difficulties so I think 0.99 from single model is very achievable. Have fun single model racing guys :D</p>",
      "rawMarkdown": "CV: 0.9913\nLB: 0.9848\n\nthough I still haven't reached a single-model score as high as @haqishen but I haven't used any TTA and big boys like seresnext or high efficientnet/densenet due to free time and computational difficulties so I think 0.99 from single model is very achievable. Have fun single model racing guys :D",
      "votes": 28,
      "replies": [
        {
          "id": 723659,
          "postDate": "2020-01-20T09:34:31.647Z",
          "content": "<p>Oh, you used magic.</p>",
          "rawMarkdown": "Oh, you used magic.",
          "votes": 3
        }
      ]
    },
    {
      "id": 742587,
      "postDate": "2020-02-11T11:52:46.237Z",
      "content": "<p>model: se-resnext50\nimg_size: 3x137x236\naugmentation: rotate, cutmix \nCV : 0.994\nLB:  0.985</p>\n\n<p>I can not get the CV score more than 0.997 from some top kagglers. So perhaps I can get some advice and help from your guys here😃 </p>",
      "rawMarkdown": "model: se-resnext50\nimg_size: 3x137x236\naugmentation: rotate, cutmix \nCV : 0.994\nLB:  0.985\n\nI can not get the CV score more than 0.997 from some top kagglers. So perhaps I can get some advice and help from your guys here😃 ",
      "votes": 22,
      "replies": [
        {
          "id": 742607,
          "postDate": "2020-02-11T12:09:05.357Z",
          "content": "<p>Great work! How many epochs did you train for and what type of lr scheduler?</p>",
          "rawMarkdown": "Great work! How many epochs did you train for and what type of lr scheduler?"
        },
        {
          "id": 742611,
          "postDate": "2020-02-11T12:11:53.787Z",
          "content": "<p>80 epochs</p>",
          "rawMarkdown": "80 epochs",
          "votes": 2
        },
        {
          "id": 742685,
          "postDate": "2020-02-11T13:38:38.387Z",
          "content": "<p>I can't surpass CV 0.982. Really curious what the top did 😲 </p>",
          "rawMarkdown": "I can't surpass CV 0.982. Really curious what the top did 😲 ",
          "votes": 2
        },
        {
          "id": 742696,
          "postDate": "2020-02-11T13:47:04.823Z",
          "content": "<p>hah.. same here😆 </p>",
          "rawMarkdown": "hah.. same here😆 "
        },
        {
          "id": 742860,
          "postDate": "2020-02-11T15:48:07.503Z",
          "content": "<p>I have tried 300 epochs and get better lb result, not use early stopping😄 </p>",
          "rawMarkdown": "I have tried 300 epochs and get better lb result, not use early stopping😄 "
        },
        {
          "id": 742871,
          "postDate": "2020-02-11T15:59:55.160Z",
          "content": "<p>300 epochs is too long😨 </p>",
          "rawMarkdown": "300 epochs is too long😨 "
        },
        {
          "id": 742935,
          "postDate": "2020-02-11T16:44:20.500Z",
          "content": "<p>How did you guys handle the data imbalance? or did you just use many augmentations and train for longer? Personally I used weights in the loss function which give me better lb score but i cant get as high of CV score as you guys!</p>",
          "rawMarkdown": "How did you guys handle the data imbalance? or did you just use many augmentations and train for longer? Personally I used weights in the loss function which give me better lb score but i cant get as high of CV score as you guys!",
          "votes": 1
        },
        {
          "id": 743019,
          "postDate": "2020-02-11T17:55:33.600Z",
          "content": "<p><a href=\"/garybios\">@garybios</a> May I wonder, did you alter standard se-resnext architecture in any way? What loss did you use? </p>",
          "rawMarkdown": "@garybios May I wonder, did you alter standard se-resnext architecture in any way? What loss did you use? ",
          "votes": 2
        },
        {
          "id": 743504,
          "postDate": "2020-02-12T05:42:47.713Z",
          "rawMarkdown": ""
        },
        {
          "id": 745944,
          "postDate": "2020-02-14T12:15:08.527Z",
          "content": "<p>Thanks for your sharing! Many similar things with you! Could you share which optimizer and LR do you choose?</p>",
          "rawMarkdown": "Thanks for your sharing! Many similar things with you! Could you share which optimizer and LR do you choose?"
        },
        {
          "id": 750495,
          "postDate": "2020-02-19T12:49:37.197Z",
          "content": "<p>May I ask what ratio do you use to balance the loss between root, vowel and consonant, if this applies to you? Thanks!</p>",
          "rawMarkdown": "May I ask what ratio do you use to balance the loss between root, vowel and consonant, if this applies to you? Thanks!"
        },
        {
          "id": 754968,
          "postDate": "2020-02-24T09:30:10.940Z",
          "content": "<p>Hi <a href=\"/garybios\">@garybios</a> , \nCan you give us an idea  about what is minimum GPU memory and run time required to train this kind of Model.</p>",
          "rawMarkdown": "Hi @garybios , \nCan you give us an idea  about what is minimum GPU memory and run time required to train this kind of Model."
        },
        {
          "id": 759248,
          "postDate": "2020-02-28T19:30:54.057Z",
          "content": "<p>Did you use a pre-trained model?</p>",
          "rawMarkdown": "Did you use a pre-trained model?"
        }
      ]
    },
    {
      "id": 734490,
      "postDate": "2020-02-01T14:56:00.117Z",
      "content": "<p>I met kernel error, so I couldn't submit :(\n```\n- img_size: 137x236\n- augmentation: auto augment, augmix\n- OHEM</p>\n\n<p>cv: 0.997\nLB: 0.9846\n```</p>",
      "rawMarkdown": "I met kernel error, so I couldn't submit :(\n```\n- img_size: 137x236\n- augmentation: auto augment, augmix\n- OHEM\n\ncv: 0.997\nLB: 0.9846\n```",
      "votes": 20,
      "replies": [
        {
          "id": 734498,
          "postDate": "2020-02-01T15:19:40.937Z",
          "content": "<p>for people like me who didn't know about OHEM; it is Online Hard Example Mining and <a href=\"http://www.erogol.com/online-hard-example-mining-pytorch/\">here</a> is a blog explaining more about it</p>",
          "rawMarkdown": "for people like me who didn't know about OHEM; it is Online Hard Example Mining and [here](http://www.erogol.com/online-hard-example-mining-pytorch/) is a blog explaining more about it",
          "votes": 13
        },
        {
          "id": 734535,
          "postDate": "2020-02-01T16:15:26.313Z",
          "content": "<p><a href=\"/bibek777\">@bibek777</a>, I was wondering what was OHEM; was it just an expression or some kind of approach. But now I'm seeing your reply here. So, Thanks. 😅 </p>",
          "rawMarkdown": "@bibek777, I was wondering what was OHEM; was it just an expression or some kind of approach. But now I'm seeing your reply here. So, Thanks. 😅 ",
          "votes": 1
        },
        {
          "id": 734539,
          "postDate": "2020-02-01T16:23:29.690Z",
          "content": "<p>thats impressive</p>",
          "rawMarkdown": "thats impressive",
          "votes": 1
        },
        {
          "id": 734592,
          "postDate": "2020-02-01T17:37:48.300Z",
          "content": "<p><a href=\"/phalanx\">@phalanx</a></p>\n\n<p>please check my starter kit, it should help you to make a submission kernel.\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757</a></p>\n\n<p>(i think you are having memory issues)</p>\n\n<p>\"cv: 0.997\", this should give you lb score of about 0.99+</p>\n\n<p>\"OHEM\" : i think this is the contribution of your score, it also confirm my (and probably other kagglers) observation:</p>\n\n<ol>\n<li>results can varies with different split</li>\n<li>some people find mixup/cutout works, other are still struggling to make it work</li>\n<li>there are only a few confusing class , most of them actually work work very well and score almost 100% in validation. further top-2 accuracy is almost 100%</li>\n<li>this competition seems to be data augmentation/sampling/synthesis problem</li>\n</ol>",
          "rawMarkdown": "@phalanx\n\nplease check my starter kit, it should help you to make a submission kernel.\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123757\n\n(i think you are having memory issues)\n\n\"cv: 0.997\", this should give you lb score of about 0.99+\n\n\"OHEM\" : i think this is the contribution of your score, it also confirm my (and probably other kagglers) observation:\n\n1. results can varies with different split\n2. some people find mixup/cutout works, other are still struggling to make it work\n3. there are only a few confusing class , most of them actually work work very well and score almost 100% in validation. further top-2 accuracy is almost 100%\n4. this competition seems to be data augmentation/sampling/synthesis problem",
          "votes": 8
        },
        {
          "id": 734625,
          "postDate": "2020-02-01T18:22:53.290Z",
          "content": "<p>Im very unfamiliar with OHEM, so my question is, does it kind of solve the imbalance dataset problem? Since it gives more attention to harder examples/bad performing examples. This might be a dumb question, i'm still a noob haha</p>",
          "rawMarkdown": "Im very unfamiliar with OHEM, so my question is, does it kind of solve the imbalance dataset problem? Since it gives more attention to harder examples/bad performing examples. This might be a dumb question, i'm still a noob haha",
          "votes": 1
        },
        {
          "id": 734708,
          "postDate": "2020-02-01T21:26:03.740Z",
          "content": "<p>Thanks. Very interesting. <br>\nI suppose if OHEM is effective, Focal-Loss should be also (I do not have any idea which is better).  IMO, OHEM will be a kind of hard-encoded Focal-Loss (top-k selected).</p>",
          "rawMarkdown": "Thanks. Very interesting.  \nI suppose if OHEM is effective, Focal-Loss should be also (I do not have any idea which is better).  IMO, OHEM will be a kind of hard-encoded Focal-Loss (top-k selected).",
          "votes": 1
        },
        {
          "id": 734774,
          "postDate": "2020-02-02T01:43:59.233Z",
          "content": "<p>mixup/cutmix with ohem loss : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
          "rawMarkdown": "mixup/cutmix with ohem loss : https://www.kaggle.com/c/bengaliai-cv19/discussion/128637",
          "votes": 9
        },
        {
          "id": 734804,
          "postDate": "2020-02-02T02:52:16.967Z",
          "content": "<p>I believe this is the original paper <a href=\"https://arxiv.org/pdf/1910.00762.pdf\">https://arxiv.org/pdf/1910.00762.pdf</a>\none advantage on focusing on top % of  biggest <code>loosers</code> is you can accelerate training =) </p>",
          "rawMarkdown": "I believe this is the original paper https://arxiv.org/pdf/1910.00762.pdf\none advantage on focusing on top % of  biggest `loosers` is you can accelerate training =) ",
          "votes": 4
        },
        {
          "id": 734829,
          "postDate": "2020-02-02T04:05:49.047Z",
          "content": "<p>is there a pytorch version of auto augment with the RL policy selection? I think it would take a long time to train</p>",
          "rawMarkdown": "is there a pytorch version of auto augment with the RL policy selection? I think it would take a long time to train"
        },
        {
          "id": 734852,
          "postDate": "2020-02-02T05:07:45.907Z",
          "content": "<p>Faster AutoAugment\n<a href=\"https://arxiv.org/abs/1911.06987\">https://arxiv.org/abs/1911.06987</a>\nExisting method use black box search algorithm, while they introduce approximate gradients to update policy and it achieve faster policy search.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2F8f918c82779294a92858014a2204d43d%2F2020-02-02%2014.06.38.png?generation=1580620035468950&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Faster AutoAugment\nhttps://arxiv.org/abs/1911.06987\nExisting method use black box search algorithm, while they introduce approximate gradients to update policy and it achieve faster policy search.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2F8f918c82779294a92858014a2204d43d%2F2020-02-02%2014.06.38.png?generation=1580620035468950&amp;alt=media)\n",
          "votes": 7
        },
        {
          "id": 734856,
          "postDate": "2020-02-02T05:16:41.323Z",
          "content": "<p>About OHEM, there are better methods than it.\nIn my experience, below 2 methods are effective for detection and classification task.\nPlease check it.</p>\n\n<p>REDUCED FOCAL LOSS\n<a href=\"https://arxiv.org/abs/1903.01347\">https://arxiv.org/abs/1903.01347</a></p>\n\n<p>Class-Balanced Loss\n<a href=\"https://arxiv.org/abs/1901.05555\">https://arxiv.org/abs/1901.05555</a></p>",
          "rawMarkdown": "About OHEM, there are better methods than it.\nIn my experience, below 2 methods are effective for detection and classification task.\nPlease check it.\n\nREDUCED FOCAL LOSS\nhttps://arxiv.org/abs/1903.01347\n\nClass-Balanced Loss\nhttps://arxiv.org/abs/1901.05555",
          "votes": 11
        },
        {
          "id": 734862,
          "postDate": "2020-02-02T05:28:24.197Z",
          "content": "<p>0.997 is a high CV for a single model.\nI'll give a try to your ideas ;)\nthanks for sharing!</p>",
          "rawMarkdown": "0.997 is a high CV for a single model.\nI'll give a try to your ideas ;)\nthanks for sharing!",
          "votes": 1
        },
        {
          "id": 734938,
          "postDate": "2020-02-02T09:01:00.883Z",
          "content": "<p>thanks for sharing.</p>",
          "rawMarkdown": "thanks for sharing.",
          "votes": 2
        },
        {
          "id": 735504,
          "postDate": "2020-02-03T05:58:19.860Z",
          "content": "<p>from my experience, 997 local can give you ~988 - 989 lb</p>",
          "rawMarkdown": "from my experience, 997 local can give you ~988 - 989 lb",
          "votes": 2
        },
        {
          "id": 735528,
          "postDate": "2020-02-03T06:46:36.263Z",
          "content": "<p>In my experiment, intensive focal loss did not work. But Reduced Focal Loss, less intensive, looks good. We can adjust a loss curve with a threshold ( cut-off) and gamma.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F2a68cc68848650a74f541083fc1f2d5d%2Ffocal-loss.png?generation=1580711991399029&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "In my experiment, intensive focal loss did not work. But Reduced Focal Loss, less intensive, looks good. We can adjust a loss curve with a threshold ( cut-off) and gamma.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F2a68cc68848650a74f541083fc1f2d5d%2Ffocal-loss.png?generation=1580711991399029&amp;alt=media)\n",
          "votes": 9
        },
        {
          "id": 736828,
          "postDate": "2020-02-04T15:56:06.013Z",
          "content": "<p>did not try OHEM yet but i dont have much luck with focal loss / reduced focal loss =(</p>",
          "rawMarkdown": "did not try OHEM yet but i dont have much luck with focal loss / reduced focal loss =(",
          "votes": 1
        },
        {
          "id": 736866,
          "postDate": "2020-02-04T16:41:50.153Z",
          "content": "<p><a href=\"/moewie94\">@moewie94</a> Same thing for me, I've tried different gamma and wasn't better than my original set up</p>",
          "rawMarkdown": "@moewie94 Same thing for me, I've tried different gamma and wasn't better than my original set up",
          "votes": 1
        },
        {
          "id": 737066,
          "postDate": "2020-02-04T21:43:52.227Z",
          "content": "<p>Thanks a lot <a href=\"/bibek777\">@bibek777</a> for such a simple explanation.</p>",
          "rawMarkdown": "Thanks a lot @bibek777 for such a simple explanation."
        }
      ]
    },
    {
      "id": 756554,
      "postDate": "2020-02-25T20:42:40.280Z",
      "content": "<p>Model: Densenet121\nimg_size: 3x224x224 (simple resize)\nAugmentation: NOT cutmix or mixup\nEpoch: 40 (still running) \nCV: 0.9938\nLB: 0.9825</p>\n\n<p>To my surprise, not all augmentation methods suit all architectures. You might have to find out which augmentation will work best for your model.</p>",
      "rawMarkdown": "Model: Densenet121\nimg_size: 3x224x224 (simple resize)\nAugmentation: NOT cutmix or mixup\nEpoch: 40 (still running) \nCV: 0.9938\nLB: 0.9825\n\nTo my surprise, not all augmentation methods suit all architectures. You might have to find out which augmentation will work best for your model.",
      "votes": 18,
      "replies": [
        {
          "id": 756572,
          "postDate": "2020-02-25T21:09:52.530Z",
          "rawMarkdown": ""
        },
        {
          "id": 756709,
          "postDate": "2020-02-26T01:38:48.360Z",
          "content": "<p><a href=\"/udaykamal\">@udaykamal</a> My guess is augmix? I recall that phalanx had a similar CV/LB. Shanks for sharing :P Really cool results</p>",
          "rawMarkdown": "@udaykamal My guess is augmix? I recall that phalanx had a similar CV/LB. Shanks for sharing :P Really cool results"
        },
        {
          "id": 757008,
          "postDate": "2020-02-26T10:35:21.800Z",
          "content": "<p>\"To my surprise, not all augmentation methods suit all architectures. You might have to find out which augmentation will work best for your model.\"</p>\n\n<p>my experiment shows that even even for the same model, different augmentation may be required for different input image size or different training epoch. (I hope i am not over fitting my validation set)</p>\n\n<p>this prompts me to look for automatic automatic augmentation hyperparameters tuning methods</p>",
          "rawMarkdown": "\"To my surprise, not all augmentation methods suit all architectures. You might have to find out which augmentation will work best for your model.\"\n\nmy experiment shows that even even for the same model, different augmentation may be required for different input image size or different training epoch. (I hope i am not over fitting my validation set)\n\nthis prompts me to look for automatic automatic augmentation hyperparameters tuning methods",
          "votes": 1
        },
        {
          "id": 757361,
          "postDate": "2020-02-26T17:16:38.637Z",
          "content": "<p>Thanks for sharing. I think it's time to get a larger model with some insights from all experiments.</p>",
          "rawMarkdown": "Thanks for sharing. I think it's time to get a larger model with some insights from all experiments."
        },
        {
          "id": 763293,
          "postDate": "2020-03-04T10:28:02.610Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 733514,
      "postDate": "2020-01-31T08:43:57.423Z",
      "content": "<p>cv: 0.9890\nlb: 0.9794</p>\n\n<p><code>\nmodel: se_resnext50_32x4d\nimgsize: 128x128\nsplit: 5/6 train, 1/6 valid\ninference: 15 minutes (kaggle kernels)\nno tta, no ensemble\n</code></p>",
      "rawMarkdown": "cv: 0.9890\nlb: 0.9794\n\n```\nmodel: se_resnext50_32x4d\nimgsize: 128x128\nsplit: 5/6 train, 1/6 valid\ninference: 15 minutes (kaggle kernels)\nno tta, no ensemble\n```\n",
      "votes": 18,
      "replies": [
        {
          "id": 733538,
          "postDate": "2020-01-31T09:21:47.050Z",
          "rawMarkdown": "",
          "votes": -5
        },
        {
          "id": 733962,
          "postDate": "2020-01-31T18:37:47.127Z",
          "rawMarkdown": "",
          "votes": 3,
          "isDeleted": true
        },
        {
          "id": 733967,
          "postDate": "2020-01-31T18:52:36.683Z",
          "content": "<p><code>tta</code> means test-time-augmentation. <code>tta</code> is a popular technique used in Kaggle to increase model performance. An example of tta would be to apply horizontal flips to test images and make prediction on flipped image, and then average with the usual image. Below is a pseudo-code for <code>tta</code>. Hope this helps you to understand the concept of <code>tta</code></p>\n\n<p>```python\nimage = get_image(img_name)\npredict_1 = model(image)</p>\n\n<p>flip_image = apply_hflip(image)\npredict_2 = model(flip_image)</p>\n\n<p>final_pred = take_mean(predict_1, predict_2)\n```</p>",
          "rawMarkdown": "`tta` means test-time-augmentation. `tta` is a popular technique used in Kaggle to increase model performance. An example of tta would be to apply horizontal flips to test images and make prediction on flipped image, and then average with the usual image. Below is a pseudo-code for `tta`. Hope this helps you to understand the concept of `tta`\n\n```python\nimage = get_image(img_name)\npredict_1 = model(image)\n\nflip_image = apply_hflip(image)\npredict_2 = model(flip_image)\n\nfinal_pred = take_mean(predict_1, predict_2)\n```\n\n",
          "votes": 16
        },
        {
          "id": 734003,
          "postDate": "2020-01-31T19:59:32.023Z",
          "content": "<p>Nice work! How many epochs did you train for?</p>",
          "rawMarkdown": "Nice work! How many epochs did you train for?",
          "votes": 1
        },
        {
          "id": 734194,
          "postDate": "2020-02-01T04:22:46.143Z",
          "content": "<p>40 epochs.</p>",
          "rawMarkdown": "40 epochs.",
          "votes": 3
        },
        {
          "id": 734263,
          "postDate": "2020-02-01T07:17:52.067Z",
          "content": "<p>Thank you for your information!\nDid you use imagenet weight ?</p>",
          "rawMarkdown": "Thank you for your information!\nDid you use imagenet weight ?",
          "votes": 1
        },
        {
          "id": 734270,
          "postDate": "2020-02-01T07:29:13.013Z",
          "content": "<p>Hi kaeru-san, \nYes, I use imagenet weights from pytorch <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">pretrainedmodels</a>. I haven't tried it without imagenet weights because I believe it should work better than random weights at least.</p>",
          "rawMarkdown": "Hi kaeru-san, \nYes, I use imagenet weights from pytorch [pretrainedmodels](https://github.com/Cadene/pretrained-models.pytorch). I haven't tried it without imagenet weights because I believe it should work better than random weights at least.",
          "votes": 2
        },
        {
          "id": 734413,
          "postDate": "2020-02-01T12:47:33.480Z",
          "content": "<p>Hi <a href=\"/appian\">@appian</a> wonderful results! I am nowhere near that with single model. If you don't mind disclosing, what is the batch size you used for training? Number of epochs is not an effective parameter without the batch size. Thx in advance</p>",
          "rawMarkdown": "Hi @appian wonderful results! I am nowhere near that with single model. If you don't mind disclosing, what is the batch size you used for training? Number of epochs is not an effective parameter without the batch size. Thx in advance"
        },
        {
          "id": 734488,
          "postDate": "2020-02-01T14:48:36.790Z",
          "content": "<p>I'm using the same model and found that without pretrained weights it performs way worse</p>",
          "rawMarkdown": "I'm using the same model and found that without pretrained weights it performs way worse"
        },
        {
          "id": 735966,
          "postDate": "2020-02-03T17:11:23.760Z",
          "content": "<p><a href=\"/appian\">@appian</a> do you use any particular trick for convergence? It's been very hard for me to reach 0.97 on validation set even after ~50 epochs.</p>",
          "rawMarkdown": "@appian do you use any particular trick for convergence? It's been very hard for me to reach 0.97 on validation set even after ~50 epochs.",
          "votes": 1
        },
        {
          "id": 737097,
          "postDate": "2020-02-04T23:06:54.993Z",
          "content": "<p><a href=\"/tahsin\">@tahsin</a> \nNot really. I just use Adam with reducelronplateau. The score is achievable with ideas discussed in this competition so far. Qishen has summarized these ideas and I think it's helpful.\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127976\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127976</a></p>",
          "rawMarkdown": "@tahsin \nNot really. I just use Adam with reducelronplateau. The score is achievable with ideas discussed in this competition so far. Qishen has summarized these ideas and I think it's helpful.\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/127976\n",
          "votes": 5
        },
        {
          "id": 737825,
          "postDate": "2020-02-05T20:19:37.780Z",
          "content": "<p><a href=\"/appian\">@appian</a> For reducelronplateau do you monitor val loss or CV?</p>",
          "rawMarkdown": "@appian For reducelronplateau do you monitor val loss or CV?",
          "votes": 1
        },
        {
          "id": 741179,
          "postDate": "2020-02-10T09:49:47.130Z",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>\nI monitored val loss with patience of 5. </p>",
          "rawMarkdown": "@greatgamedota\nI monitored val loss with patience of 5. ",
          "votes": 1
        },
        {
          "id": 741205,
          "postDate": "2020-02-10T10:15:15.917Z",
          "content": "<p><a href=\"/appian\">@appian</a> \nwould please inform what was the lowest validation loss you have got? I also monitor validation loss, my lowest validation loss was around 0.21 something. </p>",
          "rawMarkdown": "@appian \nwould please inform what was the lowest validation loss you have got? I also monitor validation loss, my lowest validation loss was around 0.21 something. ",
          "votes": 1
        },
        {
          "id": 742328,
          "postDate": "2020-02-11T08:36:08.560Z",
          "content": "<p><a href=\"/appian\">@appian</a> \nThank you for your replying !!\n(I forgot my asking you ... very sorry m(_ _)m)</p>\n\n<p>Next experiments I'll use imagenet weight !!</p>",
          "rawMarkdown": "@appian \nThank you for your replying !!\n(I forgot my asking you ... very sorry m(_ _)m)\n\nNext experiments I'll use imagenet weight !!",
          "votes": 1
        },
        {
          "id": 743912,
          "postDate": "2020-02-12T11:47:56.603Z",
          "content": "<p><a href=\"/kaerunantoka\">@kaerunantoka</a>\nNo worries. Good luck with experiments!</p>",
          "rawMarkdown": "@kaerunantoka\nNo worries. Good luck with experiments!",
          "votes": 1
        },
        {
          "id": 746084,
          "postDate": "2020-02-14T15:49:38.210Z",
          "content": "<p>Hi <a href=\"/appian\">@appian</a>, what augmentation techniques are you using? Mixup, Cutmix, Gridmask, augmix? or combination of them ?  </p>",
          "rawMarkdown": "Hi @appian, what augmentation techniques are you using? Mixup, Cutmix, Gridmask, augmix? or combination of them ?  ",
          "votes": 2
        },
        {
          "id": 751641,
          "postDate": "2020-02-20T11:10:43.277Z",
          "content": "<p>UPDATE</p>\n\n<p>CV: 0.9982\nLB: 0.9885\nepoch: 70</p>\n\n<p>single model, no tta, no ensemble. </p>",
          "rawMarkdown": "UPDATE\n\nCV: 0.9982\nLB: 0.9885\nepoch: 70\n\nsingle model, no tta, no ensemble. ",
          "votes": 11
        },
        {
          "id": 751762,
          "postDate": "2020-02-20T13:47:36.507Z",
          "content": "<p>Dude, you're killing it. Congrats</p>",
          "rawMarkdown": "Dude, you're killing it. Congrats",
          "votes": 1
        },
        {
          "id": 751781,
          "postDate": "2020-02-20T14:05:13.420Z",
          "content": "<p>May I ask if you use mixup/cutmix?</p>",
          "rawMarkdown": "May I ask if you use mixup/cutmix?",
          "votes": 1
        }
      ]
    },
    {
      "id": 762424,
      "postDate": "2020-03-03T13:54:58.717Z",
      "content": "<p><code>update</code>\ncv: 0.9981\nlb: 0.9900\nsingle fold, no tta\nI still haven't found <a href=\"https://www.kaggle.com/haqishen\">haqishen‘s</a> magic😑 😑 </p>",
      "rawMarkdown": "`update`\ncv: 0.9981\nlb: 0.9900\nsingle fold, no tta\nI still haven't found [haqishen‘s](https://www.kaggle.com/haqishen) magic😑 😑 ",
      "votes": 15,
      "replies": [
        {
          "id": 762431,
          "postDate": "2020-03-03T14:01:51.453Z",
          "content": "<p>single fold! awesome.</p>",
          "rawMarkdown": "single fold! awesome.",
          "votes": 1
        },
        {
          "id": 762446,
          "postDate": "2020-03-03T14:15:16.287Z",
          "content": "<p>Yes! </p>",
          "rawMarkdown": "Yes! "
        },
        {
          "id": 762448,
          "postDate": "2020-03-03T14:17:44.557Z",
          "content": "<p>wow</p>",
          "rawMarkdown": "wow"
        },
        {
          "id": 762471,
          "postDate": "2020-03-03T14:36:24.767Z",
          "content": "<blockquote>\n  <p>I still haven't found haqishen‘s magic😑 😑</p>\n</blockquote>\n\n<p>so you created your own 👍 👍 </p>",
          "rawMarkdown": "&gt; I still haven't found haqishen‘s magic😑 😑\n\nso you created your own 👍 👍 ",
          "votes": 2
        },
        {
          "id": 762693,
          "postDate": "2020-03-03T18:14:15.877Z",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4059244%2Faca408f0846c7509bb31f552dac21b53%2FCapture.PNG?generation=1583259244730559&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": " @ipythonx ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4059244%2Faca408f0846c7509bb31f552dac21b53%2FCapture.PNG?generation=1583259244730559&amp;alt=media)\n",
          "votes": 15
        },
        {
          "id": 762745,
          "postDate": "2020-03-03T19:22:29.333Z",
          "content": "<p>Gary NB!</p>",
          "rawMarkdown": "Gary NB!"
        },
        {
          "id": 769780,
          "postDate": "2020-03-12T09:05:20.870Z",
          "content": "<p>666</p>",
          "rawMarkdown": "666"
        }
      ]
    },
    {
      "id": 735506,
      "postDate": "2020-02-03T06:01:24.957Z",
      "content": "<p>local: <code>0.9970</code>\nlb: <code>0.9884</code></p>\n\n<p>single model, single fold, no tta</p>\n\n<p>i still think <code>0.99</code> is really doable but the you have to use big models/ tta and maybe some lucky with random seed because improvement seems quite random when local error is small</p>",
      "rawMarkdown": "local: `0.9970`\nlb: `0.9884`\n\nsingle model, single fold, no tta\n\ni still think `0.99` is really doable but the you have to use big models/ tta and maybe some lucky with random seed because improvement seems quite random when local error is small",
      "votes": 15,
      "replies": [
        {
          "id": 735521,
          "postDate": "2020-02-03T06:33:58.553Z",
          "content": "<p>wow, that' huge for single model. However I don't understand how it's so doable! I can't score minimum .97. Its seem some people are easily getting high score :( </p>\n\n<p>My basic setup: \n(keras)</p>\n\n<p><code>\nmodel: efficient \nimg size: 128\nsplit: 80/20\nopt: adam\naugmentation: augmix \nepoch: 20\n</code>\nImplementing almost same result, some people are getting really promising score.  :(</p>",
          "rawMarkdown": "wow, that' huge for single model. However I don't understand how it's so doable! I can't score minimum .97. Its seem some people are easily getting high score :( \n\nMy basic setup: \n(keras)\n\n```\nmodel: efficient \nimg size: 128\nsplit: 80/20\nopt: adam\naugmentation: augmix \nepoch: 20\n```\nImplementing almost same result, some people are getting really promising score.  :(",
          "votes": 1
        },
        {
          "id": 736840,
          "postDate": "2020-02-04T16:12:16.893Z",
          "content": "<p>Personally i haven't been having good results with efficient net maybe you should experiment with other architectures! :) </p>",
          "rawMarkdown": "Personally i haven't been having good results with efficient net maybe you should experiment with other architectures! :) ",
          "votes": 1
        },
        {
          "id": 737672,
          "postDate": "2020-02-05T16:28:26.187Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> are you using keras? If you don't mind, if you are, would you please inform the best outcome of your findings with efficientnet? </p>",
          "rawMarkdown": "@yannmajewski are you using keras? If you don't mind, if you are, would you please inform the best outcome of your findings with efficientnet? "
        },
        {
          "id": 737677,
          "postDate": "2020-02-05T16:32:54.900Z",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> If you are using augmentation like mixup/cutmix/cutout/augmix/etc. you have to train your model much longer. try it for 100-150 (or more) epochs. After 20 epochs one of my model's (cutmix) CV was ~0.968; after 120 epochs it was 0.9815</p>",
          "rawMarkdown": "@ipythonx If you are using augmentation like mixup/cutmix/cutout/augmix/etc. you have to train your model much longer. try it for 100-150 (or more) epochs. After 20 epochs one of my model's (cutmix) CV was ~0.968; after 120 epochs it was 0.9815",
          "votes": 11
        },
        {
          "id": 737687,
          "postDate": "2020-02-05T16:46:40.953Z",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> thank you for your tips. Actually I went for 200 epoch but gave an early stop for validation loss at epoch 30. I mean, if the validation loss didn't improve within 30 epoch, the training process would stop. And my model stopped at around epoch 31~33. I am not sure, should I increase the patience epoch size which is currently 30 or just pick the last optimized weights?</p>\n\n<p>Again, in this competition, an open secret is to use <code>cutmix-mixup</code> to get high score but It's comparatively hard for me to implement it right now, I am using Keras. I am currently using augmix-cutout-gridmask. Would you please inform, what is your individual validation score of the three target: <code>grapheme_root</code> , <code>vowel</code> and <code>consonant</code> ? I found that score almost  .98/.99 in both vowel and consonant is doable but score above .96/.97 for grapheme_root is pretty tough. </p>",
          "rawMarkdown": "@pestipeti thank you for your tips. Actually I went for 200 epoch but gave an early stop for validation loss at epoch 30. I mean, if the validation loss didn't improve within 30 epoch, the training process would stop. And my model stopped at around epoch 31~33. I am not sure, should I increase the patience epoch size which is currently 30 or just pick the last optimized weights?\n\nAgain, in this competition, an open secret is to use `cutmix-mixup` to get high score but It's comparatively hard for me to implement it right now, I am using Keras. I am currently using augmix-cutout-gridmask. Would you please inform, what is your individual validation score of the three target: `grapheme_root` , `vowel` and `consonant` ? I found that score almost  .98/.99 in both vowel and consonant is doable but score above .96/.97 for grapheme_root is pretty tough. ",
          "votes": 1
        },
        {
          "id": 737694,
          "postDate": "2020-02-05T16:53:39.300Z",
          "content": "<p>you should implement official metric for this competition and use this as as early stop. I am very confident this will improve your results. </p>",
          "rawMarkdown": "you should implement official metric for this competition and use this as as early stop. I am very confident this will improve your results. ",
          "votes": 3
        },
        {
          "id": 737704,
          "postDate": "2020-02-05T17:05:51.913Z",
          "content": "<p>Grapheme root: 0.9752, vowel: 0.9855, consonant: 0.9924\nI suggest (at least for one trial) remove early stopping. </p>\n\n<p>Here is a screenshot from my training, it might help. (open it in a new tab for higher resolution.)</p>\n\n<p><img src=\"https://albumizr.com/ia/8059658e005e0c58ed268812c1b1f7d1.jpg\" alt=\"\"></p>\n\n<p>Edit:\nI use iterations instead of epochs (every datapoint/validation step is after 1000 iterations/forward passes; batch: 64; valid size: 8%; ~2887 iterations = one epoch)</p>",
          "rawMarkdown": "Grapheme root: 0.9752, vowel: 0.9855, consonant: 0.9924\nI suggest (at least for one trial) remove early stopping. \n\nHere is a screenshot from my training, it might help. (open it in a new tab for higher resolution.)\n\n![](https://albumizr.com/ia/8059658e005e0c58ed268812c1b1f7d1.jpg)\n\nEdit:\nI use iterations instead of epochs (every datapoint/validation step is after 1000 iterations/forward passes; batch: 64; valid size: 8%; ~2887 iterations = one epoch)",
          "votes": 4
        },
        {
          "id": 737770,
          "postDate": "2020-02-05T18:53:35.607Z",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> Well im using a xresnet34 which i know will not give me the best results but i have computational limitations.. i added 3 tails just like <a href=\"/drhabib\">@drhabib</a> experimented in his post: \n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123432\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123432</a></p>",
          "rawMarkdown": "@ipythonx Well im using a xresnet34 which i know will not give me the best results but i have computational limitations.. i added 3 tails just like @drhabib experimented in his post: \nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123432",
          "votes": 1
        },
        {
          "id": 739024,
          "postDate": "2020-02-07T09:51:07.260Z",
          "content": "<p>Which tool are you using <a href=\"/pestipeti\">@pestipeti</a> for these visualizations? It looks amazing.</p>",
          "rawMarkdown": "Which tool are you using @pestipeti for these visualizations? It looks amazing.\n"
        },
        {
          "id": 739029,
          "postDate": "2020-02-07T10:02:35.383Z",
          "content": "<p><a href=\"/thanatoz\">@thanatoz</a> I use <a href=\"https://neptune.ai\">neptune.ai</a></p>",
          "rawMarkdown": "@thanatoz I use [neptune.ai](https://neptune.ai)",
          "votes": 1
        },
        {
          "id": 739132,
          "postDate": "2020-02-07T12:47:38.493Z",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> may I ask, when you trained a model like this \" try it for 100-150 (or more) epochs. After 20 epochs one of my model's (cutmix) CV was ~0.968; after 120 epochs it was 0.9815\", did you use any LR scheduling? </p>",
          "rawMarkdown": "@pestipeti may I ask, when you trained a model like this \" try it for 100-150 (or more) epochs. After 20 epochs one of my model's (cutmix) CV was ~0.968; after 120 epochs it was 0.9815\", did you use any LR scheduling? "
        },
        {
          "id": 739150,
          "postDate": "2020-02-07T13:22:53.533Z",
          "content": "<p><a href=\"/cateek\">@cateek</a> Yes, I used CosineAnnealingLR</p>",
          "rawMarkdown": "@cateek Yes, I used CosineAnnealingLR",
          "votes": 1
        },
        {
          "id": 741323,
          "postDate": "2020-02-10T13:35:45.003Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  whats the good learning rate and optimizer for this problem ?</p>",
          "rawMarkdown": "@drhabib  whats the good learning rate and optimizer for this problem ?",
          "votes": 1
        },
        {
          "id": 741526,
          "postDate": "2020-02-10T18:44:11.137Z",
          "content": "<p><a href=\"/mayank17\">@mayank17</a> I haven't done so much systematic studies with optimizers. So far <code>Adam</code> with LR of <code>0.003</code> and <code>ReduceOnPlatau</code> or <code>OneCycleLearning</code> policies works fine. Generally its best to stick with one optimizer and once you are satisfied with your model performance you can play around =) I will update once I have results.</p>\n\n<p>Good luck </p>",
          "rawMarkdown": "@mayank17 I haven't done so much systematic studies with optimizers. So far `Adam` with LR of `0.003` and `ReduceOnPlatau` or `OneCycleLearning` policies works fine. Generally its best to stick with one optimizer and once you are satisfied with your model performance you can play around =) I will update once I have results.\n\nGood luck ",
          "votes": 1
        },
        {
          "id": 747639,
          "postDate": "2020-02-16T17:22:57.523Z",
          "rawMarkdown": ""
        },
        {
          "id": 753198,
          "postDate": "2020-02-21T20:47:14.013Z",
          "content": "<p>mäh</p>",
          "rawMarkdown": "mäh",
          "votes": -3
        },
        {
          "id": 757578,
          "postDate": "2020-02-26T23:33:15.720Z",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Hi Peter May I ask if you pre-process the image, i.e. centered and cropped or just resize the raw image.</p>",
          "rawMarkdown": "@pestipeti Hi Peter May I ask if you pre-process the image, i.e. centered and cropped or just resize the raw image."
        },
        {
          "id": 757593,
          "postDate": "2020-02-27T00:00:43.193Z",
          "content": "<p><a href=\"/yl1202\">@yl1202</a>, for my experiment above I preprocessed the image (I used <a href=\"/iafoss\">@iafoss</a> kernel). For my current best (0.9799) I only resized the images to 128x128px. </p>",
          "rawMarkdown": "@yl1202, for my experiment above I preprocessed the image (I used @iafoss kernel). For my current best (0.9799) I only resized the images to 128x128px. "
        },
        {
          "id": 757605,
          "postDate": "2020-02-27T00:14:32.253Z",
          "content": "<p>Thanks Peter. This is really interesting...</p>",
          "rawMarkdown": "Thanks Peter. This is really interesting..."
        }
      ]
    },
    {
      "id": 761883,
      "postDate": "2020-03-03T02:28:32.213Z",
      "content": "<p>You guys may believe or not. \nYesterday, I dreamed about a solution of the top team. Even I could not remember the details, however, when I waked up, I had an idea to follow up. It gave me a major improvement 😂  </p>\n\n<p>img_size: 137x236\nCV: 0.993 \nLB: 0.9868 \nSingle fold, no tta. </p>",
      "rawMarkdown": "You guys may believe or not. \nYesterday, I dreamed about a solution of the top team. Even I could not remember the details, however, when I waked up, I had an idea to follow up. It gave me a major improvement 😂  \n\nimg_size: 137x236\nCV: 0.993 \nLB: 0.9868 \nSingle fold, no tta. ",
      "votes": 15,
      "replies": [
        {
          "id": 761884,
          "postDate": "2020-03-03T02:29:27.670Z",
          "content": "<p>I am going to sleep now. I hope to see other top's solutions 😆 </p>",
          "rawMarkdown": "I am going to sleep now. I hope to see other top's solutions 😆 ",
          "votes": 3
        },
        {
          "id": 761890,
          "postDate": "2020-03-03T02:36:20.387Z",
          "content": "<p>we probably saw the same dream ... but i remembered a little bit more...</p>\n\n<p>single model\n<code>\ncv: 0.9975\nlb: 0.9893\n</code></p>\n\n<p>single fold .. no TTA =)</p>\n\n<p>Also credit goes to my team =) </p>",
          "rawMarkdown": "we probably saw the same dream ... but i remembered a little bit more...\n\nsingle model\n```\ncv: 0.9975\nlb: 0.9893\n```\n\nsingle fold .. no TTA =)\n\nAlso credit goes to my team =) ",
          "votes": 13
        },
        {
          "id": 761895,
          "postDate": "2020-03-03T02:45:47.340Z",
          "content": "<p>I need that dream right about now... :P</p>",
          "rawMarkdown": "I need that dream right about now... :P",
          "votes": 3
        },
        {
          "id": 761896,
          "postDate": "2020-03-03T02:45:48.017Z",
          "content": "<p>I guess the solution I dreamed about is yours. :))</p>",
          "rawMarkdown": "I guess the solution I dreamed about is yours. :))",
          "votes": 1
        },
        {
          "id": 761979,
          "postDate": "2020-03-03T05:13:33.007Z",
          "content": "<p>Couldn't wait to dream your dream😄 </p>",
          "rawMarkdown": "Couldn't wait to dream your dream😄 "
        },
        {
          "id": 762006,
          "postDate": "2020-03-03T05:35:52.713Z",
          "content": "<p>Also credit goes to my team =)</p>",
          "rawMarkdown": "Also credit goes to my team =)",
          "votes": 2
        },
        {
          "id": 762461,
          "postDate": "2020-03-03T14:31:35.267Z",
          "content": "<p>What a sweet dream <a href=\"/backaggle\">@backaggle</a> and <a href=\"/drhabib\">@drhabib</a>!\nBTW, did you guys dream of any augmentation at that time 👀 ?</p>",
          "rawMarkdown": "What a sweet dream @backaggle and @drhabib!\nBTW, did you guys dream of any augmentation at that time 👀 ?",
          "votes": 1
        },
        {
          "id": 762470,
          "postDate": "2020-03-03T14:36:21.050Z",
          "content": "<p>yes I dont remember everything but it was something cutting and mixing .... perhaps cutmix ... ? and some images were rotated.... not sure.... everything is so cloudy... </p>",
          "rawMarkdown": "yes I dont remember everything but it was something cutting and mixing .... perhaps cutmix ... ? and some images were rotated.... not sure.... everything is so cloudy... ",
          "votes": 8
        },
        {
          "id": 763268,
          "postDate": "2020-03-04T10:06:04.887Z",
          "content": "<p><a href=\"/backaggle\">@backaggle</a> few days ago i saw a dream which  was \"we trained 2 models and 1 of them was trained for grapheme roots and by combining those 2 models (1 for full data and 1 for grapheme root) gave us a huge boost,i remember full dream \"i saw we are very close to gold zone after executing this plan,then didn't give it a try ha ha ha\nyou had a sweet dream :)</p>",
          "rawMarkdown": "@backaggle few days ago i saw a dream which  was \"we trained 2 models and 1 of them was trained for grapheme roots and by combining those 2 models (1 for full data and 1 for grapheme root) gave us a huge boost,i remember full dream \"i saw we are very close to gold zone after executing this plan,then didn't give it a try ha ha ha\nyou had a sweet dream :)",
          "votes": 1
        },
        {
          "id": 763687,
          "postDate": "2020-03-04T18:37:17.063Z",
          "content": "<p>I think it's time for some Inception... @christophernolan</p>",
          "rawMarkdown": "I think it's time for some Inception... @christophernolan",
          "votes": 2
        }
      ]
    },
    {
      "id": 738975,
      "postDate": "2020-02-07T08:35:04.517Z",
      "content": "<p>Model: se-resnext50-32x4d\nImage size: 128x128x1\nCV: 0.9937\nLB: 0.9838</p>\n\n<p>Interesting that no matter how I changed model structure, augmentation or image size,\nLB scores are always equal to my CV scores minus about 1~1.3%, \nguess I need totally different way to break through 99%</p>",
      "rawMarkdown": "Model: se-resnext50-32x4d\nImage size: 128x128x1\nCV: 0.9937\nLB: 0.9838\n\nInteresting that no matter how I changed model structure, augmentation or image size,\nLB scores are always equal to my CV scores minus about 1~1.3%, \nguess I need totally different way to break through 99%",
      "votes": 13,
      "replies": [
        {
          "id": 738980,
          "postDate": "2020-02-07T08:40:54.857Z",
          "content": "<p><a href=\"/ccchang801023\">@ccchang801023</a> , how many epochs are you training for?</p>",
          "rawMarkdown": "@ccchang801023 , how many epochs are you training for?"
        },
        {
          "id": 738981,
          "postDate": "2020-02-07T08:40:59.203Z",
          "content": "<p>Would you please inform what is your individual score of the three target output? </p>",
          "rawMarkdown": "Would you please inform what is your individual score of the three target output? "
        },
        {
          "id": 739020,
          "postDate": "2020-02-07T09:42:28.423Z",
          "content": "<p>And you are achieving this without Mixup/Cutmix?</p>",
          "rawMarkdown": "And you are achieving this without Mixup/Cutmix?"
        },
        {
          "id": 739233,
          "postDate": "2020-02-07T15:28:36.203Z",
          "content": "<p><a href=\"/pheadrus\">@pheadrus</a> 150 epochs\n<a href=\"/ipythonx\">@ipythonx</a>  My CV :  0.5 * 0.99107(root) + 0.25 * 0.99648(vowel) + 0.25 * 0.99641(consonant)  = 0.9937\n<a href=\"/thanatoz\">@thanatoz</a>  Yes I combined augmentation methods and Cutmix is one of them</p>",
          "rawMarkdown": "@pheadrus 150 epochs\n@ipythonx  My CV :  0.5 * 0.99107(root) + 0.25 * 0.99648(vowel) + 0.25 * 0.99641(consonant)  = 0.9937\n@thanatoz  Yes I combined augmentation methods and Cutmix is one of them",
          "votes": 7
        },
        {
          "id": 739272,
          "postDate": "2020-02-07T16:18:25.830Z",
          "content": "<p><a href=\"/ccchang801023\">@ccchang801023</a> are those co-efficient (.5, .25, .25) loss weights? </p>",
          "rawMarkdown": "@ccchang801023 are those co-efficient (.5, .25, .25) loss weights? "
        },
        {
          "id": 739283,
          "postDate": "2020-02-07T16:26:32.023Z",
          "content": "<p>No, just the weights described in evaluation metric : final_score = np.average(scores, weights=[2,1,1])</p>",
          "rawMarkdown": "No, just the weights described in evaluation metric : final_score = np.average(scores, weights=[2,1,1])",
          "votes": 2
        },
        {
          "id": 739285,
          "postDate": "2020-02-07T16:36:54.043Z",
          "content": "<p>I must say score of root is pretty impressive of yours. Have you addressed any data imbalance for it particularly or just lots of augmentation process and hard training with care? </p>",
          "rawMarkdown": "I must say score of root is pretty impressive of yours. Have you addressed any data imbalance for it particularly or just lots of augmentation process and hard training with care? ",
          "votes": 1
        },
        {
          "id": 745943,
          "postDate": "2020-02-14T12:13:18.277Z",
          "content": "<p>Which optimizer and LR do you choose?</p>",
          "rawMarkdown": "Which optimizer and LR do you choose?"
        }
      ]
    },
    {
      "id": 737631,
      "postDate": "2020-02-05T15:50:09.627Z",
      "content": "<p><strong>CV : 0.979</strong>\n*<em>LB : 0.971</em>*\n<code>\nModel: SEResNeXT50\nAugmentations: 40% Mixup + 40% Cutmix + 10% Cutout + 10% GridMask\nSplit: Character stratified split 80/20\nImage size: 128x128x3 (just resize)\nEpoch: 100\nSingle fold, No TTA\n</code></p>\n\n<p>[Next]\n- AugMix\n- Reduced Focal Loss\n- Bigger image size\n...</p>\n\n<p>There are lot of things to be done before getting to 0.99...</p>\n\n<p>I would appreciate if you could give me any advice!!</p>",
      "rawMarkdown": "**CV : 0.979**\n**LB : 0.971**\n```\nModel: SEResNeXT50\nAugmentations: 40% Mixup + 40% Cutmix + 10% Cutout + 10% GridMask\nSplit: Character stratified split 80/20\nImage size: 128x128x3 (just resize)\nEpoch: 100\nSingle fold, No TTA\n```\n\n[Next]\n- AugMix\n- Reduced Focal Loss\n- Bigger image size\n...\n\nThere are lot of things to be done before getting to 0.99...\n\nI would appreciate if you could give me any advice!!",
      "votes": 14,
      "replies": [
        {
          "id": 737636,
          "postDate": "2020-02-05T15:56:12.580Z",
          "content": "<p>Appreciate you posting details. Quick question Character stratified split is equivalent to iterative stratified splits? (<a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a>)</p>",
          "rawMarkdown": "Appreciate you posting details. Quick question Character stratified split is equivalent to iterative stratified splits? (https://github.com/trent-b/iterative-stratification)"
        },
        {
          "id": 737670,
          "postDate": "2020-02-05T16:25:15.550Z",
          "content": "<p>Impressive. Few query:\n- have you used <a href=\"/iafoss\">@iafoss</a> 's processing method?\n- epoch: 100; have you used early stop or just full 100 epoch and pick up the last updated weights?\n- about the augmentation split, would you please inform how to achieve this splitting method, I mean the percentage?\n- choosing stratified split, (its a general question) why you went only single fold? In cross validation, if you made 5 fold: 1 for test, 4 for train and you train the model. 5 fold will give you 5 times different trained model, won't it? Would you please clarify? </p>",
          "rawMarkdown": "Impressive. Few query:\n- have you used @iafoss 's processing method?\n- epoch: 100; have you used early stop or just full 100 epoch and pick up the last updated weights?\n- about the augmentation split, would you please inform how to achieve this splitting method, I mean the percentage?\n- choosing stratified split, (its a general question) why you went only single fold? In cross validation, if you made 5 fold: 1 for test, 4 for train and you train the model. 5 fold will give you 5 times different trained model, won't it? Would you please clarify? "
        },
        {
          "id": 738721,
          "postDate": "2020-02-06T22:54:53.467Z",
          "rawMarkdown": ""
        },
        {
          "id": 738989,
          "postDate": "2020-02-07T08:57:52.757Z",
          "content": "<p>Sorry for the late reply..</p>\n\n<p><a href=\"/pheadrus\">@pheadrus</a> \nCharacter stratified split may be different from  iterative stratified splits. I'm using stratified split by 1295 graphems (Bengali characters). This is because at first I have tried 1295 class classification model,  or custom models like a following figure. However, these 1295 class models were tend to overfit than 3 outputs models... Further experimentation is needed. My LB / PB score above is just a 3 outputs model.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1057275%2Fd267cf55d28be8a24e39f3ceb2534def%2F_cut.png?generation=1581063579323056&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"/ipythonx\">@ipythonx</a> </p>\n\n<blockquote>\n  <p>have you used <a href=\"/iafoss\">@iafoss</a> 's processing method?</p>\n</blockquote>\n\n<p>I haven't used it yet. I'll try it!!</p>\n\n<blockquote>\n  <p>epoch: 100; have you used early stop or just full 100 epoch and pick up the last updated weights?</p>\n</blockquote>\n\n<p>I didn't use early stopping. Further speaking, I think, 100 epochs are still not enough.</p>\n\n<blockquote>\n  <p>about the augmentation split, would you please inform how to achieve this splitting method, I mean the percentage?</p>\n</blockquote>\n\n<p>I experimented mixup only training, cutmix only training, and so on. I decided the ratio by the result, that is, mixup and cutmix were equally good, and gridmask and cutout were a little worse than mixup and cutmix. About the ratio, I need to experiment more.</p>\n\n<blockquote>\n  <p>choosing stratified split, (its a general question) why you went only single fold? In cross validation, if you made 5 fold: 1 for test, 4 for train and you train the model. 5 fold will give you 5 times different trained model, won't it? Would you please clarify?</p>\n</blockquote>\n\n<p>You are right. But, even training the model once takes a very long time. So, after I find the better method, I will finally train the model 5 times as you said. </p>\n\n<p>(Let me apologize for my poor English. )</p>",
          "rawMarkdown": "Sorry for the late reply..\n\n@pheadrus \nCharacter stratified split may be different from  iterative stratified splits. I'm using stratified split by 1295 graphems (Bengali characters). This is because at first I have tried 1295 class classification model,  or custom models like a following figure. However, these 1295 class models were tend to overfit than 3 outputs models... Further experimentation is needed. My LB / PB score above is just a 3 outputs model.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1057275%2Fd267cf55d28be8a24e39f3ceb2534def%2F_cut.png?generation=1581063579323056&amp;alt=media)\n\n@ipythonx \n\n&gt; have you used @iafoss 's processing method?\n\nI haven't used it yet. I'll try it!!\n\n\n&gt; epoch: 100; have you used early stop or just full 100 epoch and pick up the last updated weights?\n\nI didn't use early stopping. Further speaking, I think, 100 epochs are still not enough.\n\n\n&gt; about the augmentation split, would you please inform how to achieve this splitting method, I mean the percentage?\n\nI experimented mixup only training, cutmix only training, and so on. I decided the ratio by the result, that is, mixup and cutmix were equally good, and gridmask and cutout were a little worse than mixup and cutmix. About the ratio, I need to experiment more.\n\n\n&gt; choosing stratified split, (its a general question) why you went only single fold? In cross validation, if you made 5 fold: 1 for test, 4 for train and you train the model. 5 fold will give you 5 times different trained model, won't it? Would you please clarify?\n\nYou are right. But, even training the model once takes a very long time. So, after I find the better method, I will finally train the model 5 times as you said. \n\n(Let me apologize for my poor English. )",
          "votes": 10
        },
        {
          "id": 739256,
          "postDate": "2020-02-07T16:02:32.180Z",
          "content": "<p>How are you weighing the 4 losses? I think its interesting to also consider 1295 class classification. </p>",
          "rawMarkdown": "How are you weighing the 4 losses? I think its interesting to also consider 1295 class classification. "
        },
        {
          "id": 740277,
          "postDate": "2020-02-09T06:53:46.673Z",
          "content": "<p><a href=\"/pheadrus\">@pheadrus</a> \nAs shown below. (with results)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1057275%2F30898817da60d9f7500f6869bc759211%2F.png?generation=1581231174993166&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "@pheadrus \nAs shown below. (with results)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1057275%2F30898817da60d9f7500f6869bc759211%2F.png?generation=1581231174993166&amp;alt=media)\n",
          "votes": 6
        },
        {
          "id": 744362,
          "postDate": "2020-02-12T19:10:35.607Z",
          "content": "<p>What program did you use to create those figures (if you used one)? <a href=\"/inoueu1\">@inoueu1</a> </p>",
          "rawMarkdown": "What program did you use to create those figures (if you used one)? @inoueu1 "
        },
        {
          "id": 745006,
          "postDate": "2020-02-13T11:40:39.730Z",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> \nI'm using google drawings👍 </p>",
          "rawMarkdown": "@greatgamedota \nI'm using google drawings👍 ",
          "votes": 2
        }
      ]
    },
    {
      "id": 766744,
      "postDate": "2020-03-08T16:21:20.507Z",
      "content": "<p><code>\nSingle Fold:\nCV: 0.9967\nLB: 0.9884 \n</code></p>\n\n<p>EfficientNetB4, 130x224, cutmix, OHEM </p>",
      "rawMarkdown": "```\nSingle Fold:\nCV: 0.9967\nLB: 0.9884 \n```\n\nEfficientNetB4, 130x224, cutmix, OHEM ",
      "votes": 11,
      "replies": [
        {
          "id": 766764,
          "postDate": "2020-03-08T16:51:06.747Z",
          "content": "<p>any comparison results of OHEM with focal loss?</p>",
          "rawMarkdown": "any comparison results of OHEM with focal loss?",
          "votes": 1
        },
        {
          "id": 766808,
          "postDate": "2020-03-08T19:06:29.470Z",
          "content": "<p>what is OHEM？</p>",
          "rawMarkdown": "what is OHEM？",
          "votes": 2
        },
        {
          "id": 766848,
          "postDate": "2020-03-08T20:48:04.280Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> I didn't try focal loss so can't comment on that. OHEM was working well for me, so I just stuck with it. </p>\n\n<p><a href=\"/xiaolonglee\">@xiaolonglee</a> OHEM is online hard example mining. Essentially, it only backpropagates the largest losses (tunable) in each minibatch when training, the idea being that easy examples with low loss don't really contribute to learning. See here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
          "rawMarkdown": "@hengck23 I didn't try focal loss so can't comment on that. OHEM was working well for me, so I just stuck with it. \n\n@xiaolonglee OHEM is online hard example mining. Essentially, it only backpropagates the largest losses (tunable) in each minibatch when training, the idea being that easy examples with low loss don't really contribute to learning. See here: https://www.kaggle.com/c/bengaliai-cv19/discussion/128637",
          "votes": 5
        },
        {
          "id": 766851,
          "postDate": "2020-03-08T20:58:27.780Z",
          "content": "<p>Thank you for your guidence !</p>",
          "rawMarkdown": "Thank you for your guidence !",
          "votes": 1
        },
        {
          "id": 766937,
          "postDate": "2020-03-09T00:36:05.297Z",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> \nthank you for the reply.</p>\n\n<p>i am thinking that focal loss should work, but i cannot get good results with it. this puzzles me. do you have an estimate of the contributions of each step?</p>\n\n<p>in my experiments, i am having:</p>\n\n<p>EfficientNetB4, 137x236 (no augmentation) : 0.974 at local cv\nadd basic augmentation (scale, rotate, shift) : 0.978\nadd cutmix : 0.992\nadd focal loss : 0.992</p>",
          "rawMarkdown": "@vaillant \nthank you for the reply.\n\ni am thinking that focal loss should work, but i cannot get good results with it. this puzzles me. do you have an estimate of the contributions of each step?\n\nin my experiments, i am having:\n\nEfficientNetB4, 137x236 (no augmentation) : 0.974 at local cv\nadd basic augmentation (scale, rotate, shift) : 0.978\nadd cutmix : 0.992\nadd focal loss : 0.992",
          "votes": 1
        },
        {
          "id": 766961,
          "postDate": "2020-03-09T01:30:24.447Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> </p>\n\n<p>Baseline EfficientNet-B4:  0.978\nAdd cutmix: 0.986\nAdd OHEM: 0.997</p>",
          "rawMarkdown": "@hengck23 \n\nBaseline EfficientNet-B4:  0.978\nAdd cutmix: 0.986\nAdd OHEM: 0.997\n\n",
          "votes": 2
        },
        {
          "id": 767019,
          "postDate": "2020-03-09T03:28:49.453Z",
          "content": "<p>Wow, Ohem loss gave you a big boost. We found it made our se-resnext worse. Will definitely have to try on effnet. Thanks Ian</p>",
          "rawMarkdown": "Wow, Ohem loss gave you a big boost. We found it made our se-resnext worse. Will definitely have to try on effnet. Thanks Ian"
        },
        {
          "id": 767048,
          "postDate": "2020-03-09T04:58:16.263Z",
          "content": "<p>For how many epochs did you train? Is training for 100+ epochs a big factor for your result?</p>",
          "rawMarkdown": "For how many epochs did you train? Is training for 100+ epochs a big factor for your result?"
        },
        {
          "id": 767056,
          "postDate": "2020-03-09T05:15:04.973Z",
          "content": "<p>Hi <a href=\"/vaillant\">@vaillant</a> , can u pls elaborate on when did u apply OHEM during training? Did u apply OHEM from the beginning of training itself or u switched to it in the middle of training like after 50 epochs or so? </p>",
          "rawMarkdown": "Hi @vaillant , can u pls elaborate on when did u apply OHEM during training? Did u apply OHEM from the beginning of training itself or u switched to it in the middle of training like after 50 epochs or so? "
        },
        {
          "id": 767277,
          "postDate": "2020-03-09T12:17:37.573Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> How to select EfficientNet backbone, i mean b4 or b5?</p>",
          "rawMarkdown": "@hengck23 How to select EfficientNet backbone, i mean b4 or b5?"
        },
        {
          "id": 767304,
          "postDate": "2020-03-09T13:12:21.543Z",
          "content": "<p><a href=\"/kurianbenoy\">@kurianbenoy</a> 60 epochs. I tried 120 and didn't see any difference.\n<a href=\"/virajbagal\">@virajbagal</a> I don't use it at the beginning. \n<a href=\"/shayekh\">@shayekh</a> I experiment and see what works best. I favor smaller models if the performance is essentially the same. </p>",
          "rawMarkdown": "@kurianbenoy 60 epochs. I tried 120 and didn't see any difference.\n@virajbagal I don't use it at the beginning. \n@shayekh I experiment and see what works best. I favor smaller models if the performance is essentially the same. ",
          "votes": 5
        },
        {
          "id": 767671,
          "postDate": "2020-03-10T01:28:30.690Z",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> Ohem loss gave you a good boost! Thanks for sharing.One question from my side other than cutmix are you using any augmentations? </p>",
          "rawMarkdown": "@vaillant Ohem loss gave you a good boost! Thanks for sharing.One question from my side other than cutmix are you using any augmentations? ",
          "votes": 1
        },
        {
          "id": 767687,
          "postDate": "2020-03-10T02:10:34.990Z",
          "content": "<p><a href=\"/ratan123\">@ratan123</a> Nope, cutmix only.</p>",
          "rawMarkdown": "@ratan123 Nope, cutmix only.",
          "votes": 2
        },
        {
          "id": 767776,
          "postDate": "2020-03-10T04:45:39.053Z",
          "content": "<p>\"How to select EfficientNet backbone, i mean b4 or b5?\" </p>\n\n<p>see how much computation power you have.  b3,4,5 have similar performance</p>",
          "rawMarkdown": "\"How to select EfficientNet backbone, i mean b4 or b5?\" \n\nsee how much computation power you have.  b3,4,5 have similar performance",
          "votes": 1
        },
        {
          "id": 767939,
          "postDate": "2020-03-10T09:24:34.420Z",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> \n\"Baseline EfficientNet-B4: 0.978\"</p>\n\n<p>i am now trying to repeat your experiment. This is my experiment results that you may be interested in:</p>\n\n<p>Baseline EfficientNet-B3 (no augmentation #1 ) + ohem: cv 0.988 ~ 0.990\n  - size = 128x224\n  - change first convolution stride=2 to stride=1  </p>\n\n<p>[# 1]  i ran many experiments and had used previous trained models to for initialization for new training. my learning rate started at 0.05 with batch_size=256. it is possible that the final results is due to previously models (i.e. effects of initialization). But nevertheless, the results reported for  cv 0.988 ~ 0.990 is for training no augmentation for ~20 epochs after initialization</p>",
          "rawMarkdown": "@vaillant \n\"Baseline EfficientNet-B4: 0.978\"\n\ni am now trying to repeat your experiment. This is my experiment results that you may be interested in:\n\nBaseline EfficientNet-B3 (no augmentation #1 ) + ohem: cv 0.988 ~ 0.990\n  - size = 128x224\n  - change first convolution stride=2 to stride=1  \n\n [# 1]  i ran many experiments and had used previous trained models to for initialization for new training. my learning rate started at 0.05 with batch_size=256. it is possible that the final results is due to previously models (i.e. effects of initialization). But nevertheless, the results reported for  cv 0.988 ~ 0.990 is for training no augmentation for ~20 epochs after initialization",
          "votes": 1
        },
        {
          "id": 769120,
          "postDate": "2020-03-11T14:47:34.460Z",
          "content": "<p>OHEM Loss made my CV worse. I start without using OHEM, then slowly start increasing OHEM's effects as the epochs rise. Sadly did not help me in seresnext or efficientnet =)</p>",
          "rawMarkdown": "OHEM Loss made my CV worse. I start without using OHEM, then slowly start increasing OHEM's effects as the epochs rise. Sadly did not help me in seresnext or efficientnet =)",
          "votes": 1
        },
        {
          "id": 770584,
          "postDate": "2020-03-13T05:52:47.540Z",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> While OHEM was not performed well for most of the participants, you got a significant performance boost. If you don't mind, by what ratio did you use for OHEM wrt batch size?</p>",
          "rawMarkdown": "@vaillant While OHEM was not performed well for most of the participants, you got a significant performance boost. If you don't mind, by what ratio did you use for OHEM wrt batch size?"
        },
        {
          "id": 770752,
          "postDate": "2020-03-13T11:04:49.133Z",
          "content": "<p><a href=\"/muhammedazamkhan\">@muhammedazamkhan</a> I gradually decrease the ratio from 1 to 1/8 throughout training. </p>",
          "rawMarkdown": "@muhammedazamkhan I gradually decrease the ratio from 1 to 1/8 throughout training. ",
          "votes": 3
        },
        {
          "id": 771034,
          "postDate": "2020-03-13T17:23:55.497Z",
          "rawMarkdown": ""
        },
        {
          "id": 771037,
          "postDate": "2020-03-13T17:29:22.940Z",
          "rawMarkdown": ""
        },
        {
          "id": 771433,
          "postDate": "2020-03-14T07:01:49.803Z",
          "content": "<p>CV 0.994  with ohem</p>",
          "rawMarkdown": "CV 0.994  with ohem",
          "votes": -1
        }
      ]
    },
    {
      "id": 761824,
      "postDate": "2020-03-03T00:54:56.403Z",
      "content": "<p>img_size: 137x236\ncv: 0.996\nlb: 0.9894\nsingle fold, no tta</p>",
      "rawMarkdown": "img_size: 137x236\ncv: 0.996\nlb: 0.9894\nsingle fold, no tta",
      "votes": 12,
      "replies": [
        {
          "id": 761829,
          "postDate": "2020-03-03T01:07:23.337Z",
          "content": "<p>You made quantum jump :)</p>",
          "rawMarkdown": "You made quantum jump :)",
          "votes": 2
        },
        {
          "id": 761831,
          "postDate": "2020-03-03T01:11:59.907Z",
          "content": "<p>Nice! Congrats! Care to share which model?</p>",
          "rawMarkdown": "Nice! Congrats! Care to share which model?"
        },
        {
          "id": 762021,
          "postDate": "2020-03-03T05:55:57.077Z",
          "content": "<p>good job. My cv: 0.996, but lb only get 0.982</p>",
          "rawMarkdown": "good job. My cv: 0.996, but lb only get 0.982",
          "votes": 1
        }
      ]
    },
    {
      "id": 772733,
      "postDate": "2020-03-15T20:40:11.863Z",
      "content": "<p>Finally reached cv0.9975, but it's too late...\nmodel: efficientnet-b3\nsplit: 5fold with stratified k-fold\nimage size: 224x224\naugment: gridmask, cutout</p>",
      "rawMarkdown": "Finally reached cv0.9975, but it's too late...\nmodel: efficientnet-b3\nsplit: 5fold with stratified k-fold\nimage size: 224x224\naugment: gridmask, cutout",
      "votes": 9,
      "replies": [
        {
          "id": 773819,
          "postDate": "2020-03-16T00:07:45.573Z",
          "content": "<p>Nothing is too late man ;)\nLet's enjoy our last journey Lol</p>",
          "rawMarkdown": "Nothing is too late man ;)\nLet's enjoy our last journey Lol",
          "votes": 4
        },
        {
          "id": 775269,
          "postDate": "2020-03-16T13:19:28.393Z",
          "content": "<p>The best CV I can get is 0.991...</p>",
          "rawMarkdown": "The best CV I can get is 0.991...",
          "votes": 1
        },
        {
          "id": 775281,
          "postDate": "2020-03-16T13:46:27.423Z",
          "content": "<p><a href=\"/ryunosukeishizaki\">@ryunosukeishizaki</a> Thanks, I'm still training models to ensemble!!</p>",
          "rawMarkdown": "@ryunosukeishizaki Thanks, I'm still training models to ensemble!!",
          "votes": 1
        },
        {
          "id": 775282,
          "postDate": "2020-03-16T13:49:16.190Z",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> Higher cv doesn't mean higher private lb:) By the way my score is on TPU and using your gridmask implementation! Thanks!!</p>",
          "rawMarkdown": "@xiejialun Higher cv doesn't mean higher private lb:) By the way my score is on TPU and using your gridmask implementation! Thanks!!",
          "votes": 2
        },
        {
          "id": 775287,
          "postDate": "2020-03-16T13:59:29.523Z",
          "content": "<p><a href=\"/bamps53\">@bamps53</a> No problem! Good luck in private score!</p>",
          "rawMarkdown": "@bamps53 No problem! Good luck in private score!",
          "votes": 2
        },
        {
          "id": 775331,
          "postDate": "2020-03-16T15:20:57.437Z",
          "content": "<p>how many epochs?</p>",
          "rawMarkdown": "how many epochs?",
          "votes": 1
        },
        {
          "id": 775459,
          "postDate": "2020-03-16T17:44:31.580Z",
          "content": "<p><a href=\"/vishal1310\">@vishal1310</a> \nAbout 150 epochs with reduce plateau with patience=10.\nGood luck, we only have 6 hours!!</p>",
          "rawMarkdown": "@vishal1310 \nAbout 150 epochs with reduce plateau with patience=10.\nGood luck, we only have 6 hours!!",
          "votes": -1
        },
        {
          "id": 775686,
          "postDate": "2020-03-17T00:02:37.533Z",
          "content": "<p>I can see why our private scores are bad, you guys must be using one head model like see-- shared, right?\nI did notice that using one head might narrow down the prediction categorical range of data, which means models were overfitting to training and public testing data. But I did't listen to my brain, since one head model just easier to train, haha. This was a great experience!</p>",
          "rawMarkdown": "I can see why our private scores are bad, you guys must be using one head model like see-- shared, right?\nI did notice that using one head might narrow down the prediction categorical range of data, which means models were overfitting to training and public testing data. But I did't listen to my brain, since one head model just easier to train, haha. This was a great experience!",
          "votes": 1
        },
        {
          "id": 775693,
          "postDate": "2020-03-17T00:08:21.313Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2Fbb5b7423173c34fb687a1b927bc5059d%2F2020-03-17%208.05.48.png?generation=1584403618844025&amp;alt=media\" alt=\"My scores with 3 heads output\"></p>\n\n<p>My score with 3 heads output</p>",
          "rawMarkdown": "![My scores with 3 heads output](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2Fbb5b7423173c34fb687a1b927bc5059d%2F2020-03-17%208.05.48.png?generation=1584403618844025&amp;alt=media)\n\nMy score with 3 heads output"
        },
        {
          "id": 775714,
          "postDate": "2020-03-17T00:21:47.547Z",
          "content": "<p>You're absolutely right!!!\nAt first I tried to predict with 3 head, but somehow I couldn't resolve Submission Scoring Error....</p>",
          "rawMarkdown": "You're absolutely right!!!\nAt first I tried to predict with 3 head, but somehow I couldn't resolve Submission Scoring Error...."
        },
        {
          "id": 775722,
          "postDate": "2020-03-17T00:27:19.263Z",
          "content": "<p>3 heads model need more time to train and inference, so you might need to optimize the submission notebook. I took a lots of time to optimize the submission notebook, but I went to chose the one head model in the end . Well, maybe the luck will appear in next competition.</p>",
          "rawMarkdown": "3 heads model need more time to train and inference, so you might need to optimize the submission notebook. I took a lots of time to optimize the submission notebook, but I went to chose the one head model in the end . Well, maybe the luck will appear in next competition."
        },
        {
          "id": 775958,
          "postDate": "2020-03-17T03:40:16.187Z",
          "content": "<p>I took a test on my final submitted notebook.\nTotally same parameter/augmentation/model but change one head to three heads(single fold): </p>\n\n<p>1 head:\nprivate score : 0.9088\npublic score : 0.9853</p>\n\n<p>3 heads:\nprivate score : 0.9348\npublic score : 0.9797</p>",
          "rawMarkdown": "I took a test on my final submitted notebook.\nTotally same parameter/augmentation/model but change one head to three heads(single fold): \n\n1 head:\nprivate score : 0.9088\npublic score : 0.9853\n\n3 heads:\nprivate score : 0.9348\npublic score : 0.9797"
        }
      ]
    },
    {
      "id": 721664,
      "postDate": "2020-01-17T15:32:25.553Z",
      "content": "<p>I completed initial analysis using cam class activation map. Indeed there is much overfitting. Training loss can be 0.01 and validation loss can be 0.20 to 0.13. With regularisation like mixup and cutout, the cam is vastly improved.  </p>\n\n<p>Hence to get good results, one can rely on such regularization or add hand more label signal ( e.g pixel or box annotations) to improve the cam response</p>",
      "rawMarkdown": "I completed initial analysis using cam class activation map. Indeed there is much overfitting. Training loss can be 0.01 and validation loss can be 0.20 to 0.13. With regularisation like mixup and cutout, the cam is vastly improved.  \n\nHence to get good results, one can rely on such regularization or add hand more label signal ( e.g pixel or box annotations) to improve the cam response",
      "votes": 10,
      "replies": [
        {
          "id": 721767,
          "postDate": "2020-01-17T17:12:12.487Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 716280,
      "postDate": "2020-01-11T14:03:05.757Z",
      "content": "<p>[Update]\n<code>\nmodel: se-resnext50 with mixup\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.982066\nLB:  0.9739\n</code></p>",
      "rawMarkdown": "[Update]\n```\nmodel: se-resnext50 with mixup\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.982066\nLB:  0.9739\n```",
      "votes": 10,
      "replies": [
        {
          "id": 716313,
          "postDate": "2020-01-11T14:39:20.353Z",
          "content": "<p>Great result <a href=\"/bibek777\">@bibek777</a> !\nFor how many epochs did you train? Your LB score is a result of a single fold?</p>",
          "rawMarkdown": "Great result @bibek777 !\nFor how many epochs did you train? Your LB score is a result of a single fold?",
          "votes": 1
        },
        {
          "id": 716315,
          "postDate": "2020-01-11T14:41:03.483Z",
          "content": "<p>I trained it for 60 epochs. Yes, it is a single model score</p>",
          "rawMarkdown": "I trained it for 60 epochs. Yes, it is a single model score",
          "votes": 4
        },
        {
          "id": 716402,
          "postDate": "2020-01-11T16:57:07.757Z",
          "content": "<p>the difference between CV and LB after ensemble could be about 0.005.\nSince your CV is 0.982066, you can get better results after ensemble.</p>\n\n<p>(This also means that for kagglers in the rage of LB  0.985, their CV could be in the range of 0.99?)</p>",
          "rawMarkdown": "the difference between CV and LB after ensemble could be about 0.005.\nSince your CV is 0.982066, you can get better results after ensemble.\n\n(This also means that for kagglers in the rage of LB  0.985, their CV could be in the range of 0.99?)",
          "votes": 2
        },
        {
          "id": 716424,
          "postDate": "2020-01-11T17:38:29.060Z",
          "content": "<p>Did you do mix up training like you said you were trying to figure out or just normal augmentations?</p>",
          "rawMarkdown": "Did you do mix up training like you said you were trying to figure out or just normal augmentations?"
        },
        {
          "id": 716476,
          "postDate": "2020-01-11T18:51:57.950Z",
          "content": "<p>yes, surely ensemble will give better score but how far can we get with a single model? At first, I thought <code>0.970</code> but to my surprise, so far the best is <code>0.9739</code>. I wonder if we can go to <code>0.98</code> with single model?</p>",
          "rawMarkdown": "yes, surely ensemble will give better score but how far can we get with a single model? At first, I thought `0.970` but to my surprise, so far the best is `0.9739`. I wonder if we can go to `0.98` with single model?"
        },
        {
          "id": 716477,
          "postDate": "2020-01-11T18:52:45.897Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> I added mixup(+a lot of other tricks) for my current score</p>",
          "rawMarkdown": "@yannmajewski I added mixup(+a lot of other tricks) for my current score",
          "votes": 2
        },
        {
          "id": 716486,
          "postDate": "2020-01-11T19:03:03.207Z",
          "content": "<p>Wow good job my friend, youre are a beast!</p>",
          "rawMarkdown": "Wow good job my friend, youre are a beast!"
        },
        {
          "id": 716624,
          "postDate": "2020-01-12T02:38:46.750Z",
          "content": "<p>\"I wonder if we can go to 0.98 with single model?\"</p>\n\n<p>simply change 137x236 to 224x224 using F.interpolate() at the input. you may be able to see some improvement.</p>\n\n<p>control your sampling as well, e.g. mix samples of different class , mix samples of weaker class, mix samples of same class ??? mix by root, constant or vowel?</p>",
          "rawMarkdown": "\"I wonder if we can go to 0.98 with single model?\"\n\nsimply change 137x236 to 224x224 using F.interpolate() at the input. you may be able to see some improvement.\n\ncontrol your sampling as well, e.g. mix samples of different class , mix samples of weaker class, mix samples of same class ??? mix by root, constant or vowel?\n",
          "votes": 4
        },
        {
          "id": 716702,
          "postDate": "2020-01-12T05:38:14.257Z",
          "content": "<p>Interesting result. What is your loss function?</p>",
          "rawMarkdown": "Interesting result. What is your loss function?"
        },
        {
          "id": 716772,
          "postDate": "2020-01-12T08:26:43.600Z",
          "content": "<p>It is based on <a href=\"https://github.com/hysts/pytorch_image_classification/blob/master/augmentations/mixup.py\">this repo</a> </p>",
          "rawMarkdown": "It is based on [this repo](https://github.com/hysts/pytorch_image_classification/blob/master/augmentations/mixup.py) ",
          "votes": 6
        },
        {
          "id": 716782,
          "postDate": "2020-01-12T08:36:03.680Z",
          "content": "<p>60 epochs for models that deep?!\nFor me, even after optimizing the data loader as much as I can, it's gonna take ~30 hours. </p>",
          "rawMarkdown": "60 epochs for models that deep?!\nFor me, even after optimizing the data loader as much as I can, it's gonna take ~30 hours. "
        },
        {
          "id": 716952,
          "postDate": "2020-01-12T14:02:52.077Z",
          "content": "<p>Therefore, you need to unite in a team of 5 people. And train the network in parallel. (each on his own fold, for example)</p>",
          "rawMarkdown": "Therefore, you need to unite in a team of 5 people. And train the network in parallel. (each on his own fold, for example)",
          "votes": 1
        },
        {
          "id": 717716,
          "postDate": "2020-01-13T14:13:37.717Z",
          "content": "<p><a href=\"/bibek777\">@bibek777</a>  congratulation on becoming discussion Master =) </p>",
          "rawMarkdown": "@bibek777  congratulation on becoming discussion Master =) \n",
          "votes": 1
        },
        {
          "id": 720546,
          "postDate": "2020-01-16T14:15:50.747Z",
          "content": "<p>What GPU do you have?</p>",
          "rawMarkdown": "What GPU do you have?\n"
        },
        {
          "id": 720550,
          "postDate": "2020-01-16T14:16:42.113Z",
          "content": "<p>TitanX</p>",
          "rawMarkdown": "TitanX"
        }
      ]
    },
    {
      "id": 703644,
      "postDate": "2019-12-26T12:23:29.420Z",
      "content": "<p>model: ResNet 18 (pretrained=False)\nimg: 1 channel </p>\n\n<p>CV: 0.9654\nLB:  0.9632</p>\n\n<p>I think CV and LB are correlated well, according to some posts in this thread :-)</p>\n\n<hr>\n\n<p>updated on Jan.02 2020</p>\n\n<p>ResNet 18\nCV: 0.9729\nLB: 0.9681</p>",
      "rawMarkdown": "model: ResNet 18 (pretrained=False)\nimg: 1 channel \n\nCV: 0.9654\nLB:  0.9632\n\nI think CV and LB are correlated well, according to some posts in this thread :-)\n\n---\n\nupdated on Jan.02 2020\n\nResNet 18\nCV: 0.9729\nLB: 0.9681",
      "votes": 10,
      "replies": [
        {
          "id": 703668,
          "postDate": "2019-12-26T13:25:54.943Z",
          "content": "<p>Great. Few question though;</p>\n\n<ul>\n<li>what other models you've tried except resnet? </li>\n<li>cv or cross validation, right? how did you do that? I mean, have you made n times fold (of the whole train set) and train the model in that way? I know about cross validation process, but I am not sure why or how some people are doing that? Isn't it expensive for training?</li>\n<li>like others, the output are the tree labels..and using softmax for each label independently you score the probabilities, right? Have you tried to implement sigmoid on the last layer for three target variables at a time?</li>\n</ul>\n\n<p>Thank you.</p>",
          "rawMarkdown": "Great. Few question though;\n\n- what other models you've tried except resnet? \n- cv or cross validation, right? how did you do that? I mean, have you made n times fold (of the whole train set) and train the model in that way? I know about cross validation process, but I am not sure why or how some people are doing that? Isn't it expensive for training?\n- like others, the output are the tree labels..and using softmax for each label independently you score the probabilities, right? Have you tried to implement sigmoid on the last layer for three target variables at a time?\n\nThank you.\n\n\n\n",
          "votes": 4
        },
        {
          "id": 703690,
          "postDate": "2019-12-26T14:02:08.387Z",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> </p>\n\n<blockquote>\n  <p>what other models you've tried except resnet?</p>\n</blockquote>\n\n<p>Not yet now. I have just made a baseline model :-)  </p>\n\n<hr>\n\n<blockquote>\n  <p>cv or cross validation, right? how did you do that? I mean, have you made n times fold (of the whole train set) and train the model in that way? I know about cross validation process, but I am not sure why or how some people are doing that? Isn't it expensive for training?</p>\n</blockquote>\n\n<p>I used the <code>CV</code> in the meaning of Cross Validation with 5-fold. As you mentioned, cross validation process is expensive. But it will make our model evaluation robust, because we use all train data to evaluate with less leakage (Unfortunately early stopping and other hyper parameter setting can not be completely free of leakage). The process is following,</p>\n\n<p>step1. imagine you split all train data into  5 subsets(We will call these as 1, 2, 3, 4, 5) <br>\nstep2. train your model (ResNet, EfficientNet, others) with 1+2+3+4, and make inference on 5\nstep3. Do the same process as step2 (train with 2+3+4+5 and infer on 5, ...) \nstep4. Then you can get OOF(Out Of Fold) prediction on all train data and the score based on it, that's what I called CV score.</p>\n\n<hr>\n\n<blockquote>\n  <p>like others, the output are the tree labels..and using softmax for each label independently you score the probabilities, right? Have you tried to implement sigmoid on the last layer for three target variables at a time?</p>\n</blockquote>\n\n<p>Yes. I used softmax. But using sigmoid and other special structure and loss function will be alternatives.  We must do some experiments :-)</p>",
          "rawMarkdown": "@ipythonx \n&gt; what other models you've tried except resnet?\n\nNot yet now. I have just made a baseline model :-)  \n\n--- \n&gt; cv or cross validation, right? how did you do that? I mean, have you made n times fold (of the whole train set) and train the model in that way? I know about cross validation process, but I am not sure why or how some people are doing that? Isn't it expensive for training?\n\nI used the `CV` in the meaning of Cross Validation with 5-fold. As you mentioned, cross validation process is expensive. But it will make our model evaluation robust, because we use all train data to evaluate with less leakage (Unfortunately early stopping and other hyper parameter setting can not be completely free of leakage). The process is following,\n\nstep1. imagine you split all train data into  5 subsets(We will call these as 1, 2, 3, 4, 5)  \nstep2. train your model (ResNet, EfficientNet, others) with 1+2+3+4, and make inference on 5\nstep3. Do the same process as step2 (train with 2+3+4+5 and infer on 5, ...) \nstep4. Then you can get OOF(Out Of Fold) prediction on all train data and the score based on it, that's what I called CV score.\n\n---\n\n&gt;like others, the output are the tree labels..and using softmax for each label independently you score the probabilities, right? Have you tried to implement sigmoid on the last layer for three target variables at a time?\n\nYes. I used softmax. But using sigmoid and other special structure and loss function will be alternatives.  We must do some experiments :-)",
          "votes": 6
        },
        {
          "id": 703694,
          "postDate": "2019-12-26T14:07:39.467Z",
          "content": "<p>Awesome. Thank you. :-)</p>",
          "rawMarkdown": "Awesome. Thank you. :-)",
          "votes": 1
        },
        {
          "id": 703930,
          "postDate": "2019-12-26T20:33:08.480Z",
          "content": "<p>I’ve used all train set for training achieving +0.97, and LB was 0.94. Maybe overfitting.</p>",
          "rawMarkdown": "I’ve used all train set for training achieving +0.97, and LB was 0.94. Maybe overfitting.",
          "votes": 4
        },
        {
          "id": 704090,
          "postDate": "2019-12-27T03:31:53.037Z",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> How much time did your inference kernel took? \nwith 5 Fold I suspect mine will overshoot the time limit.</p>",
          "rawMarkdown": "@maxwell110 How much time did your inference kernel took? \nwith 5 Fold I suspect mine will overshoot the time limit.",
          "votes": 3
        },
        {
          "id": 704116,
          "postDate": "2019-12-27T04:28:35.970Z",
          "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> </p>\n\n<p>It will take less than 30 minutes.\nAs you know size of test data is only 12, so I simulated inference time with train data.</p>\n\n<p>One useful topic was posted, <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/122993\">https://www.kaggle.com/c/bengaliai-cv19/discussion/122993</a>\nand I think others will be.</p>\n\n<p>When we ensemble models, short inference time is desirable to blend more models. Let's take care of inference time ;-)</p>",
          "rawMarkdown": "@dhananjay3 \n\nIt will take less than 30 minutes.\nAs you know size of test data is only 12, so I simulated inference time with train data.\n\nOne useful topic was posted,  \nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/122993\nand I think others will be.\n  \nWhen we ensemble models, short inference time is desirable to blend more models. Let's take care of inference time ;-)",
          "votes": 2
        },
        {
          "id": 704190,
          "postDate": "2019-12-27T06:32:57.340Z",
          "content": "<p><a href=\"/hanjoonchoe\">@hanjoonchoe</a>  maybe too many epochs causing overfit?</p>",
          "rawMarkdown": "@hanjoonchoe  maybe too many epochs causing overfit?",
          "votes": 2
        },
        {
          "id": 704263,
          "postDate": "2019-12-27T08:27:14.257Z",
          "content": "<p><a href=\"/p4rallax\">@p4rallax</a> I don't know yet. I separated train set 8:2 now, and checking two metric score. There are a bit of gap between train and valid set as opposed to the other kagglers reporting here. Oh… Maybe They only check metrics for valid only. My bad XD. I just train my model with down sampled images so far, but using whole data set at once is more effective when I see. I do not have gpu quota to inference my model to check the score now though.</p>",
          "rawMarkdown": "@p4rallax I don't know yet. I separated train set 8:2 now, and checking two metric score. There are a bit of gap between train and valid set as opposed to the other kagglers reporting here. Oh… Maybe They only check metrics for valid only. My bad XD. I just train my model with down sampled images so far, but using whole data set at once is more effective when I see. I do not have gpu quota to inference my model to check the score now though.",
          "votes": 2
        },
        {
          "id": 704343,
          "postDate": "2019-12-27T11:00:22.927Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2791153%2Fe148f19e546d3c50c540ea24e079db8e%2FScreen%20Shot%202019-12-27%20at%206.00.02%20AM.png?generation=1577444419066240&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2791153%2Fe148f19e546d3c50c540ea24e079db8e%2FScreen%20Shot%202019-12-27%20at%206.00.02%20AM.png?generation=1577444419066240&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 707343,
          "postDate": "2019-12-31T17:10:09.910Z",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a>   Can u please describe ur training hyperparameters . I mean what was ur lr, decay and scheduler  ?</p>",
          "rawMarkdown": "@maxwell110   Can u please describe ur training hyperparameters . I mean what was ur lr, decay and scheduler  ?"
        },
        {
          "id": 715487,
          "postDate": "2020-01-10T15:28:09.890Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 766728,
      "postDate": "2020-03-08T16:05:32.770Z",
      "content": "<p>My score on <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">new validation setting</a> is like this:</p>\n\n<p><code>\nLocal Score: 0.9888  (fold_0: 38653 seen samples and 7578 unseen samples)\nLocal Score: 0.998~ (old split, single fold)\nPublic LB: 0.9923 (single fold)\n</code></p>\n\n<p>The score got around <code>0.01</code> dropped when there are 16.4% unseen graphemes in the validation set.\nIf anyone is interested in challenging validation with unseen graphemes, please have a look at this post ( <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">new validation setting</a> )</p>",
      "rawMarkdown": "My score on [new validation setting](https://www.kaggle.com/c/bengaliai-cv19/discussion/134434) is like this:\n\n```\nLocal Score: 0.9888  (fold_0: 38653 seen samples and 7578 unseen samples)\nLocal Score: 0.998~ (old split, single fold)\nPublic LB: 0.9923 (single fold)\n```\n\nThe score got around `0.01` dropped when there are 16.4% unseen graphemes in the validation set.\nIf anyone is interested in challenging validation with unseen graphemes, please have a look at this post ( [new validation setting](https://www.kaggle.com/c/bengaliai-cv19/discussion/134434) )",
      "votes": 8,
      "replies": [
        {
          "id": 766734,
          "postDate": "2020-03-08T16:10:47.823Z",
          "content": "<p>Thanks for reporting. Public LB is single model, single fold on new validation setting?</p>",
          "rawMarkdown": "Thanks for reporting. Public LB is single model, single fold on new validation setting?",
          "votes": 2
        },
        {
          "id": 766735,
          "postDate": "2020-03-08T16:11:04.083Z",
          "content": "<p>water post?😏 😏 <a href=\"/haqishen\">@haqishen</a> </p>",
          "rawMarkdown": "water post?😏 😏 @haqishen ",
          "votes": 1
        },
        {
          "id": 766758,
          "postDate": "2020-03-08T16:43:39.137Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> \nNo, LB score is not like that, before submitting I have to re-train my models on old validation splits (with all grapheme as seen samples) after experimenting on new validating set.\nThen with fully ensemble I got 0.9933.</p>",
          "rawMarkdown": "@philippsinger \nNo, LB score is not like that, before submitting I have to re-train my models on old validation splits (with all grapheme as seen samples) after experimenting on new validating set.\nThen with fully ensemble I got 0.9933.\n",
          "votes": 2
        },
        {
          "id": 766939,
          "postDate": "2020-03-09T00:41:53.097Z",
          "content": "<p>The king is back.</p>",
          "rawMarkdown": "The king is back."
        },
        {
          "id": 767000,
          "postDate": "2020-03-09T02:41:43.680Z",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> I don't get it. So the seen-unseen split is not a ready-replace plugin for e.g. Random Split? After training on the split we have to train again on old splits? I guess this is for the models to see as many graphemes as possible? Doesn't this introduce some leak into the finetuning stage, as the validation samples in stage-2 might occurred in training samples in stage 1. Wondering how you re-train on old validation splits🤥 </p>",
          "rawMarkdown": "@haqishen I don't get it. So the seen-unseen split is not a ready-replace plugin for e.g. Random Split? After training on the split we have to train again on old splits? I guess this is for the models to see as many graphemes as possible? Doesn't this introduce some leak into the finetuning stage, as the validation samples in stage-2 might occurred in training samples in stage 1. Wondering how you re-train on old validation splits🤥 ",
          "votes": 1
        },
        {
          "id": 767010,
          "postDate": "2020-03-09T03:09:03.473Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Sorry for the confusion, I'll try to make things clear.</p>\n\n<ul>\n<li>There are some unseen graphemes in the public test set, that's why we always have a gap between our CV and LB.</li>\n<li>The gap also indicated that, our models are not good at predicting unseen graphemes.</li>\n<li>We don't know how many unseen graphemes in the private test set, so there is a risk of shake-up.</li>\n<li>To address this problem, a good way is to build a validation split that similar to the public/private test set. That's why we'd better to do validation on [seen + unseen graphemes] set.</li>\n<li>Although I got some progress there, but finally I have no idea how to train my models to generalize to unseen graphemes as well as seen graphemes. </li>\n<li>So although it's kind of pain, I have to re-train all my models on old splits, to make them 'see' as much graphemes as I can.</li>\n</ul>\n\n<p><a href=\"/tonychenxyz\">@tonychenxyz</a> There's no king in Kaggle, we are all learning from each others. It's most important things to us.</p>",
          "rawMarkdown": "@roguekk007 Sorry for the confusion, I'll try to make things clear.\n\n* There are some unseen graphemes in the public test set, that's why we always have a gap between our CV and LB.\n* The gap also indicated that, our models are not good at predicting unseen graphemes.\n* We don't know how many unseen graphemes in the private test set, so there is a risk of shake-up.\n* To address this problem, a good way is to build a validation split that similar to the public/private test set. That's why we'd better to do validation on [seen + unseen graphemes] set.\n* Although I got some progress there, but finally I have no idea how to train my models to generalize to unseen graphemes as well as seen graphemes. \n* So although it's kind of pain, I have to re-train all my models on old splits, to make them 'see' as much graphemes as I can.\n\n@tonychenxyz There's no king in Kaggle, we are all learning from each others. It's most important things to us.",
          "votes": 11
        },
        {
          "id": 767016,
          "postDate": "2020-03-09T03:25:28.883Z",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> thanks for sharing such insight to tackle the overfitting issue on the private test set.  There are some diverse combinations in Bengali handwriting, a good amount of combination of those three targets and our writing styles. \nHowever, I was thinking to add more <code>grapheme_root</code> samples from the external data sources but I'm not sure whether it will help it or not also it's pain. <strong>BanglaLekha</strong> and <strong>Ekush</strong> data set is pretty good, I think. There are around 98K <code>grapheme_root</code> samples (50 categories, except ঁ ) in <strong>BanglaLekha</strong>. \nWhat do you think? And, I think <a href=\"/tonychenxyz\">@tonychenxyz</a> just wanted to appreciate your success in this competition. 😅 😃 </p>",
          "rawMarkdown": "@haqishen thanks for sharing such insight to tackle the overfitting issue on the private test set.  There are some diverse combinations in Bengali handwriting, a good amount of combination of those three targets and our writing styles. \nHowever, I was thinking to add more `grapheme_root` samples from the external data sources but I'm not sure whether it will help it or not also it's pain. **BanglaLekha** and **Ekush** data set is pretty good, I think. There are around 98K `grapheme_root` samples (50 categories, except ঁ ) in **BanglaLekha**. \nWhat do you think? And, I think @tonychenxyz just wanted to appreciate your success in this competition. 😅 😃 ",
          "votes": 1
        },
        {
          "id": 767017,
          "postDate": "2020-03-09T03:26:56.403Z",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> Thank you very much for the reply. Here I am observing some anomalies between CV and LB. It seems deeper models perform better on CV but degenerates on LB.</p>\n\n<p>I wonder how you are doing the retraining? How are you able to use stable validation set for retraining?\nI have an idea...We split the unseen graphemes into folds, and train 5-fold</p>",
          "rawMarkdown": "@haqishen Thank you very much for the reply. Here I am observing some anomalies between CV and LB. It seems deeper models perform better on CV but degenerates on LB.\n\nI wonder how you are doing the retraining? How are you able to use stable validation set for retraining?\nI have an idea...We split the unseen graphemes into folds, and train 5-fold",
          "votes": 1
        },
        {
          "id": 767022,
          "postDate": "2020-03-09T03:43:54.597Z",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> I'm not sure how they are organizing the label, and I don't have knowledge about Bengali as well... So I don't know how to do it.</p>\n\n<p><a href=\"/roguekk007\">@roguekk007</a> re-train here I mean training from  imagenet weight again.</p>",
          "rawMarkdown": "@ipythonx I'm not sure how they are organizing the label, and I don't have knowledge about Bengali as well... So I don't know how to do it.\n\n@roguekk007 re-train here I mean training from  imagenet weight again.",
          "votes": 1
        },
        {
          "id": 767030,
          "postDate": "2020-03-09T04:05:33.207Z",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> OK so the .9923 LB score has absolutely nothing to do with the new validation scheme...? You train from imagenet weight on the old validation scheme (presumably iterative stratification)? Sorry I was confused. So you are sharing the CV of the unseen validation scheme, which is .9888. I wonder if you have done any experiments submitting this to the LB?</p>",
          "rawMarkdown": "@haqishen OK so the .9923 LB score has absolutely nothing to do with the new validation scheme...? You train from imagenet weight on the old validation scheme (presumably iterative stratification)? Sorry I was confused. So you are sharing the CV of the unseen validation scheme, which is .9888. I wonder if you have done any experiments submitting this to the LB?"
        },
        {
          "id": 767036,
          "postDate": "2020-03-09T04:22:57.540Z",
          "content": "<p>If you add a trick, and get a same CV score on old split, how do you know it's useful or not?\nnew split can tell you that.</p>",
          "rawMarkdown": "If you add a trick, and get a same CV score on old split, how do you know it's useful or not?\nnew split can tell you that.",
          "votes": 1
        },
        {
          "id": 767044,
          "postDate": "2020-03-09T04:40:30.280Z",
          "content": "<p>another way to test is to use some bengali dictionary and write some unseen graphemes by hand and test (even for train)</p>\n\n<p>there are 1296 (not 1295) garphemes here:\n<a href=\"https://github.com/BengaliAI/graphemePrepare/tree/master/collection/A4\">https://github.com/BengaliAI/graphemePrepare/tree/master/collection/A4</a>\n<a href=\"https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt\">https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt</a></p>",
          "rawMarkdown": "another way to test is to use some bengali dictionary and write some unseen graphemes by hand and test (even for train)\n\nthere are 1296 (not 1295) garphemes here:\nhttps://github.com/BengaliAI/graphemePrepare/tree/master/collection/A4\nhttps://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt"
        },
        {
          "id": 767075,
          "postDate": "2020-03-09T06:04:15.943Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  It's hand written recognition competition but not hand writing competition 😂 </p>",
          "rawMarkdown": "@hengck23  It's hand written recognition competition but not hand writing competition 😂 ",
          "votes": 12
        },
        {
          "id": 767284,
          "postDate": "2020-03-09T12:41:11.530Z",
          "content": "<p>I can't read my own writing in French sometimes, I will not risk polluting my model with Bengali hand writing ;)</p>",
          "rawMarkdown": "I can't read my own writing in French sometimes, I will not risk polluting my model with Bengali hand writing ;)",
          "votes": 5
        }
      ]
    },
    {
      "id": 743999,
      "postDate": "2020-02-12T13:16:05.867Z",
      "content": "<p>model:resnet34 <br>\nimg_size: 3x 224 x 224(just resize)\naugmentation: cutmix + some augmentations\nCV: 0.985(only first fold, 80% train iterative stratified)\nLB: ???  </p>",
      "rawMarkdown": "model:resnet34  \nimg_size: 3x 224 x 224(just resize)\naugmentation: cutmix + some augmentations\nCV: 0.985(only first fold, 80% train iterative stratified)\nLB: ???  ",
      "votes": 7,
      "replies": [
        {
          "id": 750690,
          "postDate": "2020-02-19T15:58:06.393Z",
          "content": "<p>May i ask what lb score did you get with that model? I am also using resnet34, so i'm curious :)</p>",
          "rawMarkdown": "May i ask what lb score did you get with that model? I am also using resnet34, so i'm curious :)"
        },
        {
          "id": 754260,
          "postDate": "2020-02-23T09:52:58.703Z",
          "content": "<p>Sorry for the late reply. I have never submitted that. <br>\nBecause it's only experimenting for techniques  </p>\n\n<p>Now I just reached 0.988 using ResNet34</p>",
          "rawMarkdown": "Sorry for the late reply. I have never submitted that.  \nBecause it's only experimenting for techniques  \n\nNow I just reached 0.988 using ResNet34"
        },
        {
          "id": 755162,
          "postDate": "2020-02-24T14:06:32.340Z",
          "content": "<p>Great results, I'm doing pretty much exactly like you however can't get to 98%, do you think 3 channels would increase my results? thanks for your advice</p>",
          "rawMarkdown": "Great results, I'm doing pretty much exactly like you however can't get to 98%, do you think 3 channels would increase my results? thanks for your advice"
        },
        {
          "id": 755170,
          "postDate": "2020-02-24T14:12:17.983Z",
          "content": "<blockquote>\n  <p>do you think 3 channels would increase my results?</p>\n</blockquote>\n\n<p>No. I use 3 channels, but the same values among the image. <br>\npreprocess is grayscale, and I copy the same data to 3 channels.</p>",
          "rawMarkdown": "&gt; do you think 3 channels would increase my results?\n\nNo. I use 3 channels, but the same values among the image.  \npreprocess is grayscale, and I copy the same data to 3 channels."
        },
        {
          "id": 755246,
          "postDate": "2020-02-24T15:39:43.727Z",
          "content": "<p>Thanks for your advice! :)</p>",
          "rawMarkdown": "Thanks for your advice! :)"
        }
      ]
    },
    {
      "id": 731260,
      "postDate": "2020-01-28T13:14:51.337Z",
      "content": "<p>model: se-resnext50 with mixup + cutmix\nsplit: random 80/20\ninput: 3x128x128\nCV : 0.983\nLB:  0.974</p>",
      "rawMarkdown": "model: se-resnext50 with mixup + cutmix\nsplit: random 80/20\ninput: 3x128x128\nCV : 0.983\nLB:  0.974",
      "votes": 8,
      "replies": [
        {
          "id": 731448,
          "postDate": "2020-01-28T16:47:36.350Z",
          "content": "<p>Hey, what probabilities did you use for mixup and cutmix , and how many epochs did you train for? Also how did you make the image to have 3 channels?</p>",
          "rawMarkdown": "Hey, what probabilities did you use for mixup and cutmix , and how many epochs did you train for? Also how did you make the image to have 3 channels?"
        },
        {
          "id": 731733,
          "postDate": "2020-01-29T01:14:56.123Z",
          "content": "<p>（1）I trained it for 100 epochs。\n（2）img = cv2.cvtColor(img,cv2.COLOR_GRAY2BGR)。</p>",
          "rawMarkdown": "（1）I trained it for 100 epochs。\n（2）img = cv2.cvtColor(img,cv2.COLOR_GRAY2BGR)。",
          "votes": 4
        },
        {
          "id": 733479,
          "postDate": "2020-01-31T07:19:27.887Z",
          "content": "<p>Hi, did you just resize the images(137, 236) to images(128, 128)? or use something special cropping methods?</p>",
          "rawMarkdown": "Hi, did you just resize the images(137, 236) to images(128, 128)? or use something special cropping methods?",
          "votes": 1
        }
      ]
    },
    {
      "id": 770621,
      "postDate": "2020-03-13T06:58:50.963Z",
      "content": "<blockquote>\n  <p>Model: se_resnext50_32x4d\n  Image Size: 137X236\n  Train Split: 0.8, Single Fold\n  Epoch: 150\n  CV: 0.9976\n  LB: 0.9889</p>\n</blockquote>",
      "rawMarkdown": "&gt; Model: se\\_resnext50\\_32x4d\nImage Size: 137X236\nTrain Split: 0.8, Single Fold\nEpoch: 150\nCV: 0.9976\nLB: 0.9889",
      "votes": 5,
      "replies": [
        {
          "id": 770676,
          "postDate": "2020-03-13T08:45:14.527Z",
          "content": "<p>Great <a href=\"/tmheo74\">@tmheo74</a>! Can i know which augmentations are you using? </p>",
          "rawMarkdown": "Great @tmheo74! Can i know which augmentations are you using? ",
          "votes": 1
        },
        {
          "id": 770680,
          "postDate": "2020-03-13T08:53:46.757Z",
          "content": "<p>Yes, we are using cutmix, mixup, gridmask, normal aug(scale, rotate, etc.)</p>",
          "rawMarkdown": "Yes, we are using cutmix, mixup, gridmask, normal aug(scale, rotate, etc.)",
          "votes": 1
        },
        {
          "id": 770694,
          "postDate": "2020-03-13T09:25:50.970Z",
          "content": "<p>Thank you <a href=\"/tmheo74\">@tmheo74</a> All the best :)</p>",
          "rawMarkdown": "Thank you @tmheo74 All the best :)"
        },
        {
          "id": 770724,
          "postDate": "2020-03-13T10:17:59.193Z",
          "content": "<p>Man, what kind of hardware do you have? </p>",
          "rawMarkdown": "Man, what kind of hardware do you have? "
        }
      ]
    },
    {
      "id": 758464,
      "postDate": "2020-02-27T19:23:55.327Z",
      "content": "<p><code>\nmodel: seresnext50 pretrained\naug: ssr + coarse dropout w/ mixup+cutmix\ninput: 3x137x236\nsplit: iterative stratification\ntraining: adam, reducelronplateau, 23 epochs\nsingle fold CV: .97896\nLB: .9711\n</code>\nThe training isn't even done and I have a separate model with CV .985+ I haven't submitted yet as I'm waiting for it to finish training.\nIs there such thing as over sharing? Recently the big tricks have been shared and I think the LB will quickly saturate.\nEDIT:\n.985 model finished here are the results:\n<code>\nsingle fold CV: .99185\nLB: .981 (1.08 gap)\n100 epochs\nSame setup as above model but with a few changes ;)\n</code></p>",
      "rawMarkdown": "```\nmodel: seresnext50 pretrained\naug: ssr + coarse dropout w/ mixup+cutmix\ninput: 3x137x236\nsplit: iterative stratification\ntraining: adam, reducelronplateau, 23 epochs\nsingle fold CV: .97896\nLB: .9711\n```\nThe training isn't even done and I have a separate model with CV .985+ I haven't submitted yet as I'm waiting for it to finish training.\nIs there such thing as over sharing? Recently the big tricks have been shared and I think the LB will quickly saturate.\nEDIT:\n.985 model finished here are the results:\n```\nsingle fold CV: .99185\nLB: .981 (1.08 gap)\n100 epochs\nSame setup as above model but with a few changes ;)\n```",
      "votes": 5
    },
    {
      "id": 706203,
      "postDate": "2019-12-30T04:11:53.830Z",
      "content": "<p>model: B0(pretrained=True)\nimg: 1 channel \nimg_sz: original\nsplit: stratified (80/20)\noptim: AdamW\nepoch: 30\nsched: Cosine</p>\n\n<p>CV : 0.9684\nLB:  0.9629</p>\n\n<p>With similar settings, my ResNet50 also got cv 0.9633 and lb 0.9574. \nFocal loss was slightly worse than cross-entropy.</p>",
      "rawMarkdown": "model: B0(pretrained=True)\nimg: 1 channel \nimg_sz: original\nsplit: stratified (80/20)\noptim: AdamW\nepoch: 30\nsched: Cosine\n\nCV : 0.9684\nLB:  0.9629\n\nWith similar settings, my ResNet50 also got cv 0.9633 and lb 0.9574. \nFocal loss was slightly worse than cross-entropy.",
      "votes": 6,
      "replies": [
        {
          "id": 706273,
          "postDate": "2019-12-30T07:01:18.547Z",
          "content": "<p>How did you do the stratified split? </p>",
          "rawMarkdown": "How did you do the stratified split? ",
          "votes": 1
        },
        {
          "id": 707619,
          "postDate": "2020-01-01T07:05:30.043Z",
          "content": "<p><a href=\"https://www.kaggle.com/yiheng/iterative-stratification\">This</a> might help you.</p>",
          "rawMarkdown": "[This](https://www.kaggle.com/yiheng/iterative-stratification) might help you.",
          "votes": 4
        }
      ]
    },
    {
      "id": 761938,
      "postDate": "2020-03-03T03:50:29.007Z",
      "content": "<p>size: 128x128\ncv: 0.9972\nlb: 0.9890\nsingle fold, no tta</p>",
      "rawMarkdown": "size: 128x128\ncv: 0.9972\nlb: 0.9890\nsingle fold, no tta",
      "votes": 6,
      "replies": [
        {
          "id": 762199,
          "postDate": "2020-03-03T10:15:33.677Z",
          "content": "<p>Impressive for 128x128!</p>",
          "rawMarkdown": "Impressive for 128x128!",
          "votes": 1
        },
        {
          "id": 762434,
          "postDate": "2020-03-03T14:04:07.587Z",
          "content": "<p><a href=\"/nuller\">@nuller</a> hi, have you used the pre-processed image from <a href=\"https://www.kaggle.com/iafoss/grapheme-imgs-128x128\">here</a>. </p>",
          "rawMarkdown": "@nuller hi, have you used the pre-processed image from [here](https://www.kaggle.com/iafoss/grapheme-imgs-128x128). "
        },
        {
          "id": 762652,
          "postDate": "2020-03-03T17:25:09.863Z",
          "content": "<p>hi <a href=\"/ipythonx\">@ipythonx</a>, no I didn't use his preproc</p>",
          "rawMarkdown": "hi @ipythonx, no I didn't use his preproc"
        }
      ]
    },
    {
      "id": 732024,
      "postDate": "2020-01-29T11:48:55.417Z",
      "content": "<p>EfficientNet Models: [Single model , Single Fold]</p>\n\n<p>Things I tried:\nB3 - 300 x 300                                           LB: 0.9639             [40 epochs]\nB3 - 300 x 300    + cutmix                         LB: 0.9695             [40 epochs]</p>\n\n<p>B7 - 128                                                        LB : 0.9683             [40 epochs]\nB7 - 224                                                        LB : 0.9685             [15 epochs]</p>\n\n<p>Things from smaller studies:\n1.Cutmix performs better than mixup for efficientnet series\n2.Still in process of replacing Adaptiveavgpool2d with GeM\n3.Tried out multiple types of tails. Best one was similar to <a href=\"/iafoss\">@iafoss</a> \n4.Tried fade, sprinkles, cutout. Not much improvement.\n5.Using swish inplace of Mish gives me more free space in GPU. Helping me to increase the batch size</p>\n\n<p>Training longer would improve the scores. Currently planning to finalize the model. Then in later stages, to finalize parameters. Then going to run for longer epochs with multiple folds. I hope the approach I mentioned seems decent.</p>",
      "rawMarkdown": "EfficientNet Models: [Single model , Single Fold]\n\nThings I tried:\nB3 - 300 x 300                                           LB: 0.9639             [40 epochs]\nB3 - 300 x 300    + cutmix                         LB: 0.9695             [40 epochs]\n\nB7 - 128                                                        LB : 0.9683             [40 epochs]\nB7 - 224                                                        LB : 0.9685             [15 epochs]\n\nThings from smaller studies:\n1.Cutmix performs better than mixup for efficientnet series\n2.Still in process of replacing Adaptiveavgpool2d with GeM\n3.Tried out multiple types of tails. Best one was similar to @iafoss \n4.Tried fade, sprinkles, cutout. Not much improvement.\n5.Using swish inplace of Mish gives me more free space in GPU. Helping me to increase the batch size\n\nTraining longer would improve the scores. Currently planning to finalize the model. Then in later stages, to finalize parameters. Then going to run for longer epochs with multiple folds. I hope the approach I mentioned seems decent.\n\n",
      "votes": 5,
      "replies": [
        {
          "id": 733358,
          "postDate": "2020-01-31T02:30:38.850Z",
          "content": "<p>Hi! I am using efficientnet too. I wonder what LR and batch size you are using? Also, from my experiments cutmix and mixup perform best together (.5 .5 probability), maybe you should try that. Also, I wonder if you have any success implementing GeM?? For me, GeM results in training error after several epochs (floatpoint div by 0 error when using apex). I recommend you to check out the github mishcuda repository, I saw marginal improvement with mish.</p>",
          "rawMarkdown": "Hi! I am using efficientnet too. I wonder what LR and batch size you are using? Also, from my experiments cutmix and mixup perform best together (.5 .5 probability), maybe you should try that. Also, I wonder if you have any success implementing GeM?? For me, GeM results in training error after several epochs (floatpoint div by 0 error when using apex). I recommend you to check out the github mishcuda repository, I saw marginal improvement with mish."
        },
        {
          "id": 733397,
          "postDate": "2020-01-31T04:23:20.277Z",
          "content": "<p>Hey! I tried efficientnet but it seems like i can't get a good score compared to resnets. I cap at .95-.96 validation. I added 3 tails with conv/batchnorm just like my resnet models, i tried with GeM too but doesnt seem to work.. Do you have any advice?</p>",
          "rawMarkdown": "Hey! I tried efficientnet but it seems like i can't get a good score compared to resnets. I cap at .95-.96 validation. I added 3 tails with conv/batchnorm just like my resnet models, i tried with GeM too but doesnt seem to work.. Do you have any advice?"
        },
        {
          "id": 733428,
          "postDate": "2020-01-31T05:28:31.090Z",
          "content": "<p>I remodified the GeM and its working better than adpativeavgpool2d as mentioned here. I have tested for only one epoch runs. </p>",
          "rawMarkdown": "I remodified the GeM and its working better than adpativeavgpool2d as mentioned here. I have tested for only one epoch runs. "
        },
        {
          "id": 733579,
          "postDate": "2020-01-31T10:17:45.220Z",
          "content": "<p>Are you using rwightman implementation or lukemelas implementation</p>",
          "rawMarkdown": "Are you using rwightman implementation or lukemelas implementation\n\n"
        },
        {
          "id": 734213,
          "postDate": "2020-02-01T05:07:25.497Z",
          "content": "<p>Im using this implementation of GeM:</p>\n\n<p>```\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def <strong>init</strong>(self, p=3, eps=1e-6):\n        super(GeM,self).<strong>init</strong>()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def <strong>repr</strong>(self):\n        return self.<strong>class</strong>._<em>name</em>_ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'</p>\n\n<p>```\nWhich is working well on other models. Have you tried a cnn/batchnorm tail with efficientnet? I tried but it doesnt seem to do so well</p>",
          "rawMarkdown": "Im using this implementation of GeM:\n\n```\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n\n```\nWhich is working well on other models. Have you tried a cnn/batchnorm tail with efficientnet? I tried but it doesnt seem to do so well"
        },
        {
          "id": 734254,
          "postDate": "2020-02-01T07:01:22.557Z",
          "content": "<p>Also efficient, b4 with 224x224, rotate, cutmix and mixup, finally lb 0.9682 (100 epochs) </p>",
          "rawMarkdown": "Also efficient, b4 with 224x224, rotate, cutmix and mixup, finally lb 0.9682 (100 epochs) ",
          "votes": 2
        },
        {
          "id": 736441,
          "postDate": "2020-02-04T07:22:52.693Z",
          "content": "<p>Most models will get WxH in the range of 16x16 or something similar. So if you use \"Conv\" in the tails it will work. But for efficientnet. Its 2560x4x4.  So using a conv layer on the tail doesn't going to capture much information.</p>\n\n<p>So dont prefer conv layer for efficientnet. I am opening a thread specific for efficientnet ppl. So we can discuss and improve the score. </p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128911\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128911</a> </p>",
          "rawMarkdown": "Most models will get WxH in the range of 16x16 or something similar. So if you use \"Conv\" in the tails it will work. But for efficientnet. Its 2560x4x4.  So using a conv layer on the tail doesn't going to capture much information.\n\nSo dont prefer conv layer for efficientnet. I am opening a thread specific for efficientnet ppl. So we can discuss and improve the score. \n\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/128911 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 726402,
      "postDate": "2020-01-23T01:56:50.750Z",
      "content": "<p>model: xresnet34(pretrained=False)\nimage size: 128\noptim: Lookahead(SGD)\nsched: OneCycleLr\naugmentations: RandomResizedCrop, RandomRotate, RandomErasing, RandomPerspective</p>\n\n<p>CV: 0.9740\nLB: 0.9672</p>",
      "rawMarkdown": "model: xresnet34(pretrained=False)\nimage size: 128\noptim: Lookahead(SGD)\nsched: OneCycleLr\naugmentations: RandomResizedCrop, RandomRotate, RandomErasing, RandomPerspective\n\nCV: 0.9740\nLB: 0.9672",
      "votes": 5,
      "replies": [
        {
          "id": 730135,
          "postDate": "2020-01-27T06:20:23.127Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Wow! That is a very charming score considering your configuration. I was not able to reach that score without using some magic. Do you mind telling if you used some tricks for this score?</p>\n\n<p>Also, my LB and CV correlation is nearly identical as yours: ~.007 gap. Cheers</p>",
          "rawMarkdown": "@yannmajewski Wow! That is a very charming score considering your configuration. I was not able to reach that score without using some magic. Do you mind telling if you used some tricks for this score?\n\nAlso, my LB and CV correlation is nearly identical as yours: ~.007 gap. Cheers",
          "votes": 1
        },
        {
          "id": 730461,
          "postDate": "2020-01-27T14:08:26.393Z",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> What i found worked the best after many tests:\n -3 tails, one for each category\n-Adding a weight tensor to the loss function to handle imbalanced classes\n-Training for longer (60+ epochs)\n-Training with SGD and OneCycleLr gave me better results with the same setup</p>\n\n<p>Other than that just hyperparameter tuning!</p>\n\n<p>I seem to not be able to have a better result with mixup/cutmix, do you have any advice? What magic did you use? :p</p>",
          "rawMarkdown": "@roguekk007 What i found worked the best after many tests:\n -3 tails, one for each category\n-Adding a weight tensor to the loss function to handle imbalanced classes\n-Training for longer (60+ epochs)\n-Training with SGD and OneCycleLr gave me better results with the same setup\n\nOther than that just hyperparameter tuning!\n\nI seem to not be able to have a better result with mixup/cutmix, do you have any advice? What magic did you use? :p",
          "votes": 5
        },
        {
          "id": 730823,
          "postDate": "2020-01-28T01:20:07.213Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> \nI am using cutmix and mixup with .5 probability each :) Here is my implementation. There highly-voted posts out there, too. For me, I found a .0005 moderate improvement. With onecycle I also observe better scores, but I have to tune the number of epochs (which is extremely time-consuming). Still testing out how much more training cutmix+mixup requires. For me, the rule of the thumb is 1.5-2 times more training.</p>\n\n<p>Copied and edited from a wide range of sources\n```\ndef rand_bbox(size, lam):\n    # Revision because there is only 1 channel\n    W = size[1]\n    H = size[2]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)</p>\n\n<pre><code># uniform\ncx = np.random.randint(W)\ncy = np.random.randint(H)\n\nbbx1 = np.clip(cx - cut_w // 2, 0, W)\nbby1 = np.clip(cy - cut_h // 2, 0, H)\nbbx2 = np.clip(cx + cut_w // 2, 0, W)\nbby2 = np.clip(cy + cut_h // 2, 0, H)\n\nreturn bbx1, bby1, bbx2, bby2\n</code></pre>\n\n<p><code>\n</code></p>\n\n<h1>For training loop</h1>\n\n<pre><code>    x, labels1, labels2, labels3 = batch\n   # Cutmixup here\n    choice = np.random.rand(1)\n    if choice &amp;lt;= args['cutmix']:\n        # Cutmix!\n        lam = np.random.beta(1., 1.)\n        rand_index = torch.randperm(x.shape[0]).cuda()\n        l1_a, l2_a, l3_a = labels1, labels2, labels3\n        l1_b, l2_b, l3_b = l1_a[rand_index], l2_a[rand_index], l3_a[rand_index]\n        bbx1, bby1, bbx2, bby2 = rand_bbox(x.size(), lam)\n        x[:, bbx1:bbx2, bby1:bby2] = x[rand_index, bbx1:bbx2, bby1:bby2]\n        # adjust lambda to exactly match pixel ratio\n        lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (x.size()[-1] * x.size()[-2]))\n        # compute output\n        preds1, preds2, preds3 = model(x)\n        loss = criterion(preds1, preds2, preds3, l1_a, l2_a, l3_a) * lam + \\\n                criterion(preds1, preds2, preds3, l1_b, l2_b, l3_b) * (1 - lam)\n    elif choice &amp;lt;= args['cutmix'] + args['mixup']:\n        # Mixup!\n        indices = torch.randperm(x.size(0))\n        shuffled_x = x[indices]\n        shuffled_l1, shuffled_l2, shuffled_l3 = labels1[indices], labels2[indices], labels3[indices]\n        lam = np.random.beta(.4, .4)\n        x = x * lam + shuffled_x * (1 - lam)\n        preds1, preds2, preds3 = model(x)\n        loss = lam * criterion(preds1, preds2, preds3, labels1, labels2, labels3) +\\\n                        (1 - lam) * criterion(preds1, preds2, preds3, shuffled_l1, shuffled_l2, shuffled_l3)\n    else:\n        preds1, preds2, preds3 = model(x)\n        loss = criterion(preds1, preds2, preds3, labels1, labels2, labels3)\n\n    with amp.scale_loss(loss, op) as scaled_loss:\n        scaled_loss.backward()\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "@yannmajewski \nI am using cutmix and mixup with .5 probability each :) Here is my implementation. There highly-voted posts out there, too. For me, I found a .0005 moderate improvement. With onecycle I also observe better scores, but I have to tune the number of epochs (which is extremely time-consuming). Still testing out how much more training cutmix+mixup requires. For me, the rule of the thumb is 1.5-2 times more training.\n\nCopied and edited from a wide range of sources\n```\ndef rand_bbox(size, lam):\n    # Revision because there is only 1 channel\n    W = size[1]\n    H = size[2]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)\n\n    # uniform\n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // 2, 0, W)\n    bby1 = np.clip(cy - cut_h // 2, 0, H)\n    bbx2 = np.clip(cx + cut_w // 2, 0, W)\n    bby2 = np.clip(cy + cut_h // 2, 0, H)\n\n    return bbx1, bby1, bbx2, bby2\n```\n```\n\n# For training loop\n        x, labels1, labels2, labels3 = batch\n       # Cutmixup here\n        choice = np.random.rand(1)\n        if choice &lt;= args['cutmix']:\n            # Cutmix!\n            lam = np.random.beta(1., 1.)\n            rand_index = torch.randperm(x.shape[0]).cuda()\n            l1_a, l2_a, l3_a = labels1, labels2, labels3\n            l1_b, l2_b, l3_b = l1_a[rand_index], l2_a[rand_index], l3_a[rand_index]\n            bbx1, bby1, bbx2, bby2 = rand_bbox(x.size(), lam)\n            x[:, bbx1:bbx2, bby1:bby2] = x[rand_index, bbx1:bbx2, bby1:bby2]\n            # adjust lambda to exactly match pixel ratio\n            lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (x.size()[-1] * x.size()[-2]))\n            # compute output\n            preds1, preds2, preds3 = model(x)\n            loss = criterion(preds1, preds2, preds3, l1_a, l2_a, l3_a) * lam + \\\n                    criterion(preds1, preds2, preds3, l1_b, l2_b, l3_b) * (1 - lam)\n        elif choice &lt;= args['cutmix'] + args['mixup']:\n            # Mixup!\n            indices = torch.randperm(x.size(0))\n            shuffled_x = x[indices]\n            shuffled_l1, shuffled_l2, shuffled_l3 = labels1[indices], labels2[indices], labels3[indices]\n            lam = np.random.beta(.4, .4)\n            x = x * lam + shuffled_x * (1 - lam)\n            preds1, preds2, preds3 = model(x)\n            loss = lam * criterion(preds1, preds2, preds3, labels1, labels2, labels3) +\\\n                            (1 - lam) * criterion(preds1, preds2, preds3, shuffled_l1, shuffled_l2, shuffled_l3)\n        else:\n            preds1, preds2, preds3 = model(x)\n            loss = criterion(preds1, preds2, preds3, labels1, labels2, labels3)\n\n        with amp.scale_loss(loss, op) as scaled_loss:\n            scaled_loss.backward()\n```",
          "votes": 4
        },
        {
          "id": 730833,
          "postDate": "2020-01-28T01:52:56.803Z",
          "content": "<p>Thanks! I saw that post, and tried it but it seems like i couldn't get a better score, maybe i didnt train long enough.. </p>",
          "rawMarkdown": "Thanks! I saw that post, and tried it but it seems like i couldn't get a better score, maybe i didnt train long enough.. "
        }
      ]
    },
    {
      "id": 715130,
      "postDate": "2020-01-10T06:40:17.173Z",
      "content": "<p><code>\nmodel: se-resnext50\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.977519\nLB:  0.9682\n</code></p>",
      "rawMarkdown": "```\nmodel: se-resnext50\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.977519\nLB:  0.9682\n```",
      "votes": 5,
      "replies": [
        {
          "id": 716017,
          "postDate": "2020-01-11T06:36:16.717Z",
          "content": "<p>So you didn't resize the image rather than go with raw size! Have you used any image preprocessing of your own? or data augmentation? Seeing your CV score, how many folds did you make? If you made, let's say 5 fold and got 5 models, then did you pick up the single best scoring model for submission? </p>\n\n<p>However, CV implementation is pesky for the multi-label problem. Would you like to please view this thread where I've mentioned a <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#715487\">CV issue</a>? </p>",
          "rawMarkdown": "So you didn't resize the image rather than go with raw size! Have you used any image preprocessing of your own? or data augmentation? Seeing your CV score, how many folds did you make? If you made, let's say 5 fold and got 5 models, then did you pick up the single best scoring model for submission? \n\nHowever, CV implementation is pesky for the multi-label problem. Would you like to please view this thread where I've mentioned a [CV issue](https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#715487)? "
        },
        {
          "id": 716021,
          "postDate": "2020-01-11T06:48:10.987Z",
          "content": "<p>I didn't use any image preprocessing. this is a single fold score based on 80/20 split of data. </p>",
          "rawMarkdown": "I didn't use any image preprocessing. this is a single fold score based on 80/20 split of data. "
        },
        {
          "id": 716187,
          "postDate": "2020-01-11T11:27:36.323Z",
          "rawMarkdown": "",
          "votes": -1
        },
        {
          "id": 763286,
          "postDate": "2020-03-04T10:24:36.987Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 719121,
      "postDate": "2020-01-15T06:45:32.930Z",
      "content": "<p><code>\nif np.random.rand()&amp;lt;0.5: do mixup\nelse: do cutmix\n</code></p>\n\n<p>i am thinking of</p>\n\n<p>```\nrandom.choice([mixup, cutmix, cutout, sprinkle, augmix, etc ...])</p>\n\n<p>```</p>",
      "rawMarkdown": "```\nif np.random.rand()&lt;0.5: do mixup\nelse: do cutmix\n```\n\ni am thinking of\n\n```\nrandom.choice([mixup, cutmix, cutout, sprinkle, augmix, etc ...])\n\n```",
      "votes": 6,
      "replies": [
        {
          "id": 721603,
          "postDate": "2020-01-17T14:22:53.327Z",
          "content": "<p>cutout helps. </p>\n\n<p><code>\nmodel: se-resnext50 (pertained= True)\naugment: cutout, rotate(20)\nsplit: random 80/20\ninput: 1 x 128 x 128\nValidation Score : 0.9821900\nLB:  0.9745\n</code>\nwithout cutout ruining with the same set up I was getting LB <code>0.969</code></p>\n\n<p>Image preprocessing - with <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a></p>",
          "rawMarkdown": "cutout helps. \n\n```\nmodel: se-resnext50 (pertained= True)\naugment: cutout, rotate(20)\nsplit: random 80/20\ninput: 1 x 128 x 128\nValidation Score : 0.9821900\nLB:  0.9745\n```\nwithout cutout ruining with the same set up I was getting LB `0.969`\n\nImage preprocessing - with https://www.kaggle.com/iafoss/image-preprocessing-128x128",
          "votes": 14
        },
        {
          "id": 721643,
          "postDate": "2020-01-17T15:01:26.963Z",
          "rawMarkdown": ""
        },
        {
          "id": 721644,
          "postDate": "2020-01-17T15:02:02.857Z",
          "content": "<p>Impressive. However, One thing I like clarify. As you said you split the training set 80:20, how you state validation score as a CV score. Isn’t it when we do KFold cross validation, then we say about CV score! So, is it cross validation (CV) or simply just the validation score.</p>",
          "rawMarkdown": "Impressive. However, One thing I like clarify. As you said you split the training set 80:20, how you state validation score as a CV score. Isn’t it when we do KFold cross validation, then we say about CV score! So, is it cross validation (CV) or simply just the validation score.",
          "votes": 1
        },
        {
          "id": 721646,
          "postDate": "2020-01-17T15:04:25.883Z",
          "content": "<p>Ahhh... Thanks for bringing this up. I guess I should be clear. Its a validation score not 5 fold CV =) I will correct </p>",
          "rawMarkdown": "Ahhh... Thanks for bringing this up. I guess I should be clear. Its a validation score not 5 fold CV =) I will correct ",
          "votes": 1
        },
        {
          "id": 721657,
          "postDate": "2020-01-17T15:22:59.687Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 721715,
          "postDate": "2020-01-17T16:34:04.257Z",
          "content": "<p>I cannot use image size of 128 x 128 with pretrained se-resnext models due to problems with somewhere the image is resized into say 2048 x -2 x -2 :p I found that the existing architecture indeed results in this error. So, can you give me any advice on how you do this? <a href=\"https://www.kaggle.com/drhabib\"></a><a href=\"/drhabib\">@drhabib</a></p>",
          "rawMarkdown": "I cannot use image size of 128 x 128 with pretrained se-resnext models due to problems with somewhere the image is resized into say 2048 x -2 x -2 :p I found that the existing architecture indeed results in this error. So, can you give me any advice on how you do this? [@drhabib](https://www.kaggle.com/drhabib)",
          "votes": 2
        },
        {
          "id": 721724,
          "postDate": "2020-01-17T16:42:00.323Z",
          "content": "<p>Ahh  are you using Pytorch ? </p>\n\n<p>it seems like you have to add Polling Layer and than flatten. </p>\n\n<p><code>\n1) First cut the model\n2) add  nn.AdaptiveAvgPool2d(1) \n3) Flatten ()\n4) nn.Linear\n</code></p>\n\n<p>Let me know if you need more detailed example</p>",
          "rawMarkdown": "Ahh  are you using Pytorch ? \n\nit seems like you have to add Polling Layer and than flatten. \n\n```\n1) First cut the model\n2) add  nn.AdaptiveAvgPool2d(1) \n3) Flatten ()\n4) nn.Linear\n```\n\nLet me know if you need more detailed example",
          "votes": 4
        },
        {
          "id": 721771,
          "postDate": "2020-01-17T17:18:33.930Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> thanks for your clarification. However, one thing more, the validation score; well I'm assuming you've used softmax and cross_entropy_loss. So, you should have three validation score, such as: grapheme root, vowel, and consonant. So, do you mention the average validation score of them? </p>",
          "rawMarkdown": "@drhabib thanks for your clarification. However, one thing more, the validation score; well I'm assuming you've used softmax and cross_entropy_loss. So, you should have three validation score, such as: grapheme root, vowel, and consonant. So, do you mention the average validation score of them? ",
          "votes": 1
        },
        {
          "id": 721779,
          "postDate": "2020-01-17T17:26:16.583Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Thank you very much for the explanation! I think I got it. I might bug you again if I face any problem! :p Sorry in advance.</p>",
          "rawMarkdown": "@drhabib Thank you very much for the explanation! I think I got it. I might bug you again if I face any problem! :p Sorry in advance.",
          "votes": 1
        },
        {
          "id": 721788,
          "postDate": "2020-01-17T17:38:06.220Z",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> \nOk so I just train one model with 80% being training data and 20% validation.</p>\n\n<p>scores for individual groups:</p>\n\n<p><code>\ngrapheme - 0.976273\nvowel - 0.990575\nconstant  - 0.985638\n</code></p>\n\n<p><code>\nlocal score using competition metric: 0.9821900\nlb scores: 0.9745\n</code></p>\n\n<p>I hope its clear =) </p>",
          "rawMarkdown": "@ipythonx \nOk so I just train one model with 80% being training data and 20% validation.\n\nscores for individual groups:\n\n```\ngrapheme - 0.976273\nvowel - 0.990575\nconstant  - 0.985638\n```\n\n```\nlocal score using competition metric: 0.9821900\nlb scores: 0.9745\n```\n\nI hope its clear =) ",
          "votes": 3
        },
        {
          "id": 721794,
          "postDate": "2020-01-17T17:44:08.530Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> That sure does. Thank you. 😃 </p>",
          "rawMarkdown": "@drhabib That sure does. Thank you. 😃 ",
          "votes": 1
        },
        {
          "id": 721850,
          "postDate": "2020-01-17T19:30:32.470Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> thanks, I have a few questions : what alpha do you use for cutout (for the beta distribution)?  do you perform it at every batch (or only x% of batches)? what does your training score looks like (can you still overfit)?</p>",
          "rawMarkdown": "@drhabib thanks, I have a few questions : what alpha do you use for cutout (for the beta distribution)?  do you perform it at every batch (or only x% of batches)? what does your training score looks like (can you still overfit)?"
        },
        {
          "id": 721880,
          "postDate": "2020-01-17T20:10:02.503Z",
          "content": "<p>actually one can google for \"mixup\" and\"cutout\" for more advanced versions or other variations. find papers that cite the original papers.</p>\n\n<p>e.g. here is a variation just published</p>\n\n<p>GridMask Data Augmentation\nP Chen - arXiv preprint arXiv:2001.04086, 2020 - arxiv.org\n3 days ago - We propose a novel data augmentation methodGridMask'in this paper. It\nutilizes information removal to achieve state-of-the-art results in a variety of computer vision\ntasks. We analyze the requirement of information dropping. Then we show limitation of …</p>",
          "rawMarkdown": "actually one can google for \"mixup\" and\"cutout\" for more advanced versions or other variations. find papers that cite the original papers.\n\ne.g. here is a variation just published\n\nGridMask Data Augmentation\nP Chen - arXiv preprint arXiv:2001.04086, 2020 - arxiv.org\n3 days ago - We propose a novel data augmentation methodGridMask'in this paper. It\nutilizes information removal to achieve state-of-the-art results in a variety of computer vision\ntasks. We analyze the requirement of information dropping. Then we show limitation of …",
          "votes": 5
        },
        {
          "id": 722002,
          "postDate": "2020-01-18T01:22:16.773Z",
          "content": "<p>Thanks <a href=\"/hengck23\">@hengck23</a>  . And for everyones use here is the official code for Gridmask </p>\n\n<p><a href=\"https://github.com/akuxcw/GridMask\">https://github.com/akuxcw/GridMask</a></p>",
          "rawMarkdown": "Thanks @hengck23  . And for everyones use here is the official code for Gridmask \n\nhttps://github.com/akuxcw/GridMask",
          "votes": 1
        },
        {
          "id": 722698,
          "postDate": "2020-01-18T23:58:49.310Z",
          "content": "<p>here is some random gridmask applied to the dataset <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1203323%2F08b26252532a94ecf5028b025d66bb19%2Fgridmask.png?generation=1579391912490101&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "here is some random gridmask applied to the dataset ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1203323%2F08b26252532a94ecf5028b025d66bb19%2Fgridmask.png?generation=1579391912490101&amp;alt=media)\n",
          "votes": 10
        },
        {
          "id": 723906,
          "postDate": "2020-01-20T15:37:56.140Z",
          "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a> , it seems that I am severely overfitting. Can u tell what optimizer, scheduler u used? What are ur learning rates and any weight decay ? Thanks. </p>",
          "rawMarkdown": "Hi @drhabib , it seems that I am severely overfitting. Can u tell what optimizer, scheduler u used? What are ur learning rates and any weight decay ? Thanks. ",
          "votes": 1
        },
        {
          "id": 723943,
          "postDate": "2020-01-20T16:07:46.173Z",
          "content": "<p>Hi <a href=\"/virajbagal\">@virajbagal</a> <br>\nI would say you should worry very little about optimizers, schedulers and learning rates. For example I can tell you can reach top 30 with Adam or SDG with learning rate anywhere in between (0.01-0.003). </p>\n\n<p>I would also recommend to choose one model, e.g Resnet50 or Seresnext (more bigger) and do all your experiments. </p>\n\n<p>You correctly identified the problem which is overfitting on training data. How this problem can be solved ? I recommend you go thru all the highly upvoted posts in this thread and write down there approaches. Make a list and go one by one. There is no other way.  </p>\n\n<p>After you done this, you can focus on optimizers,  hyperparamters and bigger models. </p>\n\n<p>Hope to see you in top 50 =) </p>",
          "rawMarkdown": "Hi @virajbagal  \nI would say you should worry very little about optimizers, schedulers and learning rates. For example I can tell you can reach top 30 with Adam or SDG with learning rate anywhere in between (0.01-0.003). \n\nI would also recommend to choose one model, e.g Resnet50 or Seresnext (more bigger) and do all your experiments. \n\nYou correctly identified the problem which is overfitting on training data. How this problem can be solved ? I recommend you go thru all the highly upvoted posts in this thread and write down there approaches. Make a list and go one by one. There is no other way.  \n\nAfter you done this, you can focus on optimizers,  hyperparamters and bigger models. \n\nHope to see you in top 50 =) ",
          "votes": 19
        },
        {
          "id": 723952,
          "postDate": "2020-01-20T16:27:26.077Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  i should probably print your golden suggestion .</p>",
          "rawMarkdown": "@drhabib  i should probably print your golden suggestion .",
          "votes": 2
        },
        {
          "id": 723956,
          "postDate": "2020-01-20T16:35:48.527Z",
          "content": "<p>You should be giving advice to us, since you are in top 10 =) </p>",
          "rawMarkdown": "You should be giving advice to us, since you are in top 10 =) ",
          "votes": 2
        },
        {
          "id": 724544,
          "postDate": "2020-01-21T08:28:04.873Z",
          "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a> , thank you for ur advice. I am following the top posts but I am not able to reproduce the results. It seems many of u guys have easily reached 0.96+ using only simple augs, random splits and normal learning methods,  but here I am struggling to even cross 0.9595 . I don't understand where I am going wrong. Anyways, it is all about trying out I guess. I'll try even harder xD </p>",
          "rawMarkdown": "Hi @drhabib , thank you for ur advice. I am following the top posts but I am not able to reproduce the results. It seems many of u guys have easily reached 0.96+ using only simple augs, random splits and normal learning methods,  but here I am struggling to even cross 0.9595 . I don't understand where I am going wrong. Anyways, it is all about trying out I guess. I'll try even harder xD ",
          "votes": 2
        },
        {
          "id": 724791,
          "postDate": "2020-01-21T13:42:37.170Z",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> I see the frustration =)  I would recommend then looking on top 3 scoring kernels with 0.96+ and try to tweak them. </p>",
          "rawMarkdown": "@virajbagal I see the frustration =)  I would recommend then looking on top 3 scoring kernels with 0.96+ and try to tweak them. ",
          "votes": 2
        },
        {
          "id": 726441,
          "postDate": "2020-01-23T02:26:25.053Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> when you use a pretrained model do you freeze some layers and then train? Thanks in advance!</p>",
          "rawMarkdown": "@drhabib when you use a pretrained model do you freeze some layers and then train? Thanks in advance!"
        },
        {
          "id": 727195,
          "postDate": "2020-01-23T14:21:19.360Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> it did not help so much. so I <code>unfreeze</code> everything and train.</p>\n\n<p>Good luck</p>",
          "rawMarkdown": "@yannmajewski it did not help so much. so I `unfreeze` everything and train.\n\nGood luck",
          "votes": 3
        }
      ]
    },
    {
      "id": 703378,
      "postDate": "2019-12-26T04:22:01.640Z",
      "content": "<p>mine:\n<code>\nresnet34\nimg_sz: 128\nsched: OneCycleWithWarmup\nsplit: random (80/20)\nimg: 3 channels\nepoch: 15\nCV: 0.9610\nLB: 0.9643\n</code></p>",
      "rawMarkdown": "mine:\n```\nresnet34\nimg_sz: 128\nsched: OneCycleWithWarmup\nsplit: random (80/20)\nimg: 3 channels\nepoch: 15\nCV: 0.9610\nLB: 0.9643\n```",
      "votes": 5,
      "replies": [
        {
          "id": 704683,
          "postDate": "2019-12-27T20:38:56.207Z",
          "content": "<p><a href=\"/backaggle\">@backaggle</a> - just interested to know if you extended the grayscale image to color in 3 channels. Or if you just replicated the grayscale images to extend to three channels. Also, how does this impact the learning? I would imagine this wouldn't make a large difference but I could be wrong I am not familiar with using grayscale as more than 1 channel.</p>",
          "rawMarkdown": "@backaggle - just interested to know if you extended the grayscale image to color in 3 channels. Or if you just replicated the grayscale images to extend to three channels. Also, how does this impact the learning? I would imagine this wouldn't make a large difference but I could be wrong I am not familiar with using grayscale as more than 1 channel.",
          "votes": 3
        },
        {
          "id": 706222,
          "postDate": "2019-12-30T05:07:48.973Z",
          "content": "<p><a href=\"/backaggle\">@backaggle</a>  Hi. Did u use resnet34 pretrained weights ? </p>",
          "rawMarkdown": "@backaggle  Hi. Did u use resnet34 pretrained weights ? ",
          "votes": 1
        }
      ]
    },
    {
      "id": 762152,
      "postDate": "2020-03-03T08:56:44.057Z",
      "content": "<p>CV: 0.9890\nLB: 0.9845</p>\n\n<p>Looking for +1 teammember with higher CV and LB</p>",
      "rawMarkdown": "CV: 0.9890\nLB: 0.9845\n\nLooking for +1 teammember with higher CV and LB",
      "votes": 3,
      "replies": [
        {
          "id": 763435,
          "postDate": "2020-03-04T13:41:00.803Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Haha just flipping through the discussions one last time before the end....seeing this message... :) </p>",
          "rawMarkdown": "@philippsinger Haha just flipping through the discussions one last time before the end....seeing this message... :) ",
          "votes": 2
        },
        {
          "id": 765396,
          "postDate": "2020-03-06T15:24:31.833Z",
          "content": "<p>Looks like you don't need another teammember anymore 😃 </p>",
          "rawMarkdown": "Looks like you don't need another teammember anymore 😃 ",
          "votes": 2
        }
      ]
    },
    {
      "id": 756846,
      "postDate": "2020-02-26T06:23:07.660Z",
      "content": "<p>model: se-resnext50\nimg_size: 128, chenel: 3\naugumentation: cutmix * autoaugment\nepoch:100\ncv: 0.984\nlb: 0.9786</p>\n\n<p>Around next week, I can use gpu(v100).\nSo I want to try img size 224.</p>",
      "rawMarkdown": "model: se-resnext50\nimg_size: 128, chenel: 3\naugumentation: cutmix * autoaugment\nepoch:100\ncv: 0.984\nlb: 0.9786\n\nAround next week, I can use gpu(v100).\nSo I want to try img size 224.",
      "votes": 3,
      "replies": [
        {
          "id": 757456,
          "postDate": "2020-02-26T19:28:34.260Z",
          "content": "<p><a href=\"/kani23\">@kani23</a> What is your training time with GPU for 100 epochs?</p>",
          "rawMarkdown": "@kani23 What is your training time with GPU for 100 epochs?"
        },
        {
          "id": 758157,
          "postDate": "2020-02-27T13:50:07.133Z",
          "content": "<p>about 12 hr.</p>",
          "rawMarkdown": "about 12 hr."
        }
      ]
    },
    {
      "id": 732272,
      "postDate": "2020-01-29T16:23:48.797Z",
      "content": "<p>Update 2\n<code>\nmodel: seresnext50 pretrained w/ mixup+cutmix\naffine augmentation\ninput: 3x128x128\nsplit: iterative stratification\ntraining: adam, reducelronplateau, 50-70 epochs\nCV: .97177\nLB: .9643\n</code>\nUpdate 3\n<code>\nsame model\n.4 mixup, .4 cutmix, .2 aug only\naffine aug\nsame input\nsame split\nadamw, one cycle lr, 100 epochs\nCV: .97699\nLB: .9661\n</code>\nsame tail as <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\">https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964</a></p>",
      "rawMarkdown": "Update 2\n```\nmodel: seresnext50 pretrained w/ mixup+cutmix\naffine augmentation\ninput: 3x128x128\nsplit: iterative stratification\ntraining: adam, reducelronplateau, 50-70 epochs\nCV: .97177\nLB: .9643\n```\nUpdate 3\n```\nsame model\n.4 mixup, .4 cutmix, .2 aug only\naffine aug\nsame input\nsame split\nadamw, one cycle lr, 100 epochs\nCV: .97699\nLB: .9661\n```\nsame tail as https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964",
      "votes": 3
    },
    {
      "id": 711915,
      "postDate": "2020-01-06T16:50:00.230Z",
      "content": "<p>Update:\n<code>\nmodel: se-resnext50\nAugmentations: Random rotate(+-15 degree), Zoom(1.1), Cutout etc.\nimg: 1 channel \ndim: 128\nEpoch: 25\noptim: over9000\nCV : 0.9704\nLB:  0.9651\n</code></p>",
      "rawMarkdown": "Update:\n```\nmodel: se-resnext50\nAugmentations: Random rotate(+-15 degree), Zoom(1.1), Cutout etc.\nimg: 1 channel \ndim: 128\nEpoch: 25\noptim: over9000\nCV : 0.9704\nLB:  0.9651\n```",
      "votes": 3,
      "replies": [
        {
          "id": 712856,
          "postDate": "2020-01-07T16:35:37.590Z",
          "content": "<p>Hey! I'm currently using the same model and similar augmentations, however I'm not able to get similar results. This is the first time im using the over9000 optimizer and im struggling to find good hyperparameters. Could you tell me what are your parameters for your optim? Just trying to learn, thanks!! :)</p>",
          "rawMarkdown": "Hey! I'm currently using the same model and similar augmentations, however I'm not able to get similar results. This is the first time im using the over9000 optimizer and im struggling to find good hyperparameters. Could you tell me what are your parameters for your optim? Just trying to learn, thanks!! :)",
          "votes": 1
        },
        {
          "id": 712903,
          "postDate": "2020-01-07T17:47:24.017Z",
          "content": "<p>I'm not particularly using any trick. Just trying to create a baseline. Maybe you can train your model with SGD optimizer for a long time(~30 epochs or more) and try to hit 0.97 or higher on your validation data. Let us know what happens.</p>",
          "rawMarkdown": "I'm not particularly using any trick. Just trying to create a baseline. Maybe you can train your model with SGD optimizer for a long time(~30 epochs or more) and try to hit 0.97 or higher on your validation data. Let us know what happens.",
          "votes": 1
        },
        {
          "id": 712913,
          "postDate": "2020-01-07T18:15:45.243Z",
          "content": "<p>Right now my best model was achieved with a resnet18 and adam, im getting a little more than 0.97 on my validation data but when using over9000 it doesn't seem to increase so that's why i asked you that! But thanks for your answer! :)</p>",
          "rawMarkdown": "Right now my best model was achieved with a resnet18 and adam, im getting a little more than 0.97 on my validation data but when using over9000 it doesn't seem to increase so that's why i asked you that! But thanks for your answer! :)",
          "votes": 2
        },
        {
          "id": 713294,
          "postDate": "2020-01-08T05:49:50.710Z",
          "content": "<p>How do you split your data into train-val? I am using <a href=\"https://www.kaggle.com/yiheng/iterative-stratification\">iterative stratification</a>. The gap between my CV and LB is around 0.005~0.007. So when my model scores 0.97 or higher on val set, I assume that it's gonna score &gt;0.96 on LB. Also, it's been really hard for me to reach and go beyond 0.97 using <code>over9000</code>. I think I need to do some experiments with other optimizers.</p>",
          "rawMarkdown": "How do you split your data into train-val? I am using [iterative stratification](https://www.kaggle.com/yiheng/iterative-stratification). The gap between my CV and LB is around 0.005~0.007. So when my model scores 0.97 or higher on val set, I assume that it's gonna score &gt;0.96 on LB. Also, it's been really hard for me to reach and go beyond 0.97 using `over9000`. I think I need to do some experiments with other optimizers.",
          "votes": 2
        },
        {
          "id": 713359,
          "postDate": "2020-01-08T07:40:14.990Z",
          "content": "<p><a href=\"/tahsin\">@tahsin</a> I'm using the same iterative stratification and I can relate to your drop on the leaderboard, the only way I found so far to reach my CV score on LB is to perform TTA with each fold's model and do a voting. I'd be happy to know about a better CV scheme too!</p>\n\n<p>Also, do you have random augmentation during your training/validation time? do you perform early stopping? My thoughts on why I see a drop is that my early stopping is actually randomly overfitting each validation set because of validation augmentation (early stopping will end up when random augmentation will yield a randomly good performance). Maybe I should stop doing augmentation on the validation set to avoid this problem. Will try and let you know.</p>",
          "rawMarkdown": "@tahsin I'm using the same iterative stratification and I can relate to your drop on the leaderboard, the only way I found so far to reach my CV score on LB is to perform TTA with each fold's model and do a voting. I'd be happy to know about a better CV scheme too!\n\nAlso, do you have random augmentation during your training/validation time? do you perform early stopping? My thoughts on why I see a drop is that my early stopping is actually randomly overfitting each validation set because of validation augmentation (early stopping will end up when random augmentation will yield a randomly good performance). Maybe I should stop doing augmentation on the validation set to avoid this problem. Will try and let you know.",
          "votes": 1
        },
        {
          "id": 713620,
          "postDate": "2020-01-08T13:32:04.197Z",
          "content": "<p>I'm randomly splitting 80/20. I'm not even cross validating because i'm a noob and i'm not sure how to do so, I just use simple validation curve.. I think my problem was that my root accuracy was much lower than the two others, i was able to increase it by adding a weight of the loss of each class</p>",
          "rawMarkdown": "I'm randomly splitting 80/20. I'm not even cross validating because i'm a noob and i'm not sure how to do so, I just use simple validation curve.. I think my problem was that my root accuracy was much lower than the two others, i was able to increase it by adding a weight of the loss of each class"
        },
        {
          "id": 714338,
          "postDate": "2020-01-09T10:17:51.990Z",
          "content": "<p><a href=\"/tahsin\">@tahsin</a>  Hi. I am little comfused. By 25 epochs, do u mean u train on 4 folds with 25 epochs keeping 1 fold as validation, then u train on other 4 folds  for 25 epochs keeping 1 fold as validation ,i.e, in total 125 epochs ? <strong>OR</strong> u just train for 5 epochs on  each group of 4 folds and ur total epochs are 25? </p>",
          "rawMarkdown": "@tahsin  Hi. I am little comfused. By 25 epochs, do u mean u train on 4 folds with 25 epochs keeping 1 fold as validation, then u train on other 4 folds  for 25 epochs keeping 1 fold as validation ,i.e, in total 125 epochs ? **OR** u just train for 5 epochs on  each group of 4 folds and ur total epochs are 25? "
        },
        {
          "id": 714748,
          "postDate": "2020-01-09T17:57:17.670Z",
          "content": "<p><a href=\"/optimo\">@optimo</a> I used augmentations as I mentioned before during training. And I'm not using early stopping. Please let us know the outcomes of your experiments. </p>",
          "rawMarkdown": "@optimo I used augmentations as I mentioned before during training. And I'm not using early stopping. Please let us know the outcomes of your experiments. "
        },
        {
          "id": 714752,
          "postDate": "2020-01-09T18:02:53.480Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> don't be so hard on yourself! I think for starter you are doing great. My root score is also comparatively lower than others. Maybe I will train a different classifier for root or use a <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123432\">long tail</a>.   </p>",
          "rawMarkdown": "@yannmajewski don't be so hard on yourself! I think for starter you are doing great. My root score is also comparatively lower than others. Maybe I will train a different classifier for root or use a [long tail](https://www.kaggle.com/c/bengaliai-cv19/discussion/123432).   ",
          "votes": 1
        },
        {
          "id": 714756,
          "postDate": "2020-01-09T18:06:59.657Z",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> I did not use cross-validation. I trained a single model for 25 epochs.</p>",
          "rawMarkdown": "@virajbagal I did not use cross-validation. I trained a single model for 25 epochs.",
          "votes": 1
        },
        {
          "id": 715687,
          "postDate": "2020-01-10T18:14:25.657Z",
          "content": "<p><a href=\"/tahsin\">@tahsin</a>  Hi. What do u mean by 1 channel, i.e, Are you using a new conv layer having 3 filters  at the beginning of se_resnext50 to make image of 3 channels and then passing thru the original se_resnext50 model ?</p>",
          "rawMarkdown": "@tahsin  Hi. What do u mean by 1 channel, i.e, Are you using a new conv layer having 3 filters  at the beginning of se_resnext50 to make image of 3 channels and then passing thru the original se_resnext50 model ?"
        },
        {
          "id": 715775,
          "postDate": "2020-01-10T20:01:56.677Z",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> I am modifying the first conv layer to accept 1 channel images instead of 3.</p>",
          "rawMarkdown": "@virajbagal I am modifying the first conv layer to accept 1 channel images instead of 3."
        },
        {
          "id": 715798,
          "postDate": "2020-01-10T21:12:11.863Z",
          "content": "<p><a href=\"/tahsin\">@tahsin</a> where did you get your model? I can't seem to find it</p>",
          "rawMarkdown": "@tahsin where did you get your model? I can't seem to find it"
        },
        {
          "id": 716644,
          "postDate": "2020-01-12T03:46:30.130Z",
          "content": "<p><a href=\"/tahsin\">@tahsin</a>  what learning rate scheduler are u using ? </p>",
          "rawMarkdown": "@tahsin  what learning rate scheduler are u using ? "
        },
        {
          "id": 716838,
          "postDate": "2020-01-12T10:49:32.457Z",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> I am using <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">pretrainedmodels</a> library. You can use <code>se-resnext50</code> from there. </p>",
          "rawMarkdown": "@yannmajewski I am using [pretrainedmodels](https://github.com/Cadene/pretrained-models.pytorch) library. You can use `se-resnext50` from there. ",
          "votes": 1
        },
        {
          "id": 716840,
          "postDate": "2020-01-12T10:51:03.440Z",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> I am using <code>cyclic learning rate</code></p>",
          "rawMarkdown": "@virajbagal I am using `cyclic learning rate`",
          "votes": 1
        },
        {
          "id": 717584,
          "postDate": "2020-01-13T10:27:07.393Z",
          "content": "<p><a href=\"/tahsin\">@tahsin</a>  Hi. What loss function are you using? </p>",
          "rawMarkdown": "@tahsin  Hi. What loss function are you using? "
        },
        {
          "id": 719049,
          "postDate": "2020-01-15T04:16:33.460Z",
          "content": "<p>Combined weighted CE loss. You can find it in <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\">this</a> kernel. </p>",
          "rawMarkdown": "Combined weighted CE loss. You can find it in [this](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964) kernel. "
        },
        {
          "id": 723859,
          "postDate": "2020-01-20T14:43:06.043Z",
          "content": "<p>Hi <a href=\"/tahsin\">@tahsin</a> , I think I am facing severe overfitting problems. Can you tell what are ur base_lr and max_lr for cyclic learning rate and weight decays for over9k ?  Thanks </p>",
          "rawMarkdown": "Hi @tahsin , I think I am facing severe overfitting problems. Can you tell what are ur base_lr and max_lr for cyclic learning rate and weight decays for over9k ?  Thanks "
        }
      ]
    },
    {
      "id": 734858,
      "postDate": "2020-02-02T05:22:49.393Z",
      "content": "<p>I would like to add this as comparison to ohem <a href=\"https://github.com/qychen13/DifficultyAwareEmbedding\">https://github.com/qychen13/DifficultyAwareEmbedding</a></p>",
      "rawMarkdown": "I would like to add this as comparison to ohem https://github.com/qychen13/DifficultyAwareEmbedding",
      "votes": 4,
      "replies": [
        {
          "id": 735331,
          "postDate": "2020-02-02T22:12:45.803Z",
          "content": "<p>this also seems to be interesting work:\n<a href=\"https://arxiv.org/pdf/1912.11188v1.pdf\">https://arxiv.org/pdf/1912.11188v1.pdf</a>\nADVERSARIAL AUTOAUGMENT (ICLR 2020)</p>\n\n<p>\" The augmentation policy network attempts to increase the training loss of a target network through generating adversarial augmentation policies, while the target network can learn more robust features from harder examples to improve the generalization. \"</p>",
          "rawMarkdown": "this also seems to be interesting work:\nhttps://arxiv.org/pdf/1912.11188v1.pdf\nADVERSARIAL AUTOAUGMENT (ICLR 2020)\n\n\" The augmentation policy network attempts to increase the training loss of a target network through generating adversarial augmentation policies, while the target network can learn more robust features from harder examples to improve the generalization. \"",
          "votes": 2
        }
      ]
    },
    {
      "id": 705279,
      "postDate": "2019-12-28T17:46:42.823Z",
      "content": "<p><a href=\"/drhabib\">@drhabib</a>   Have u used any augmentations ? </p>",
      "rawMarkdown": "@drhabib   Have u used any augmentations ? ",
      "votes": 3,
      "replies": [
        {
          "id": 705410,
          "postDate": "2019-12-28T22:40:47.857Z",
          "content": "<p>Hey there =) For this particular experiment I have used just one <code>random_rotate(10)</code></p>",
          "rawMarkdown": "Hey there =) For this particular experiment I have used just one `random_rotate(10)`",
          "votes": 5
        },
        {
          "id": 705818,
          "postDate": "2019-12-29T14:20:25.077Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Is your current result using a single model? </p>",
          "rawMarkdown": "@drhabib Is your current result using a single model? ",
          "votes": 2
        },
        {
          "id": 705823,
          "postDate": "2019-12-29T14:29:59.477Z",
          "content": "<p>I am using B0 efficientet. And 5 fold CV results in my current standing. </p>",
          "rawMarkdown": "I am using B0 efficientet. And 5 fold CV results in my current standing. ",
          "votes": 2
        },
        {
          "id": 705826,
          "postDate": "2019-12-29T14:40:48.417Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  Hi. Are u training locally ?  If no, then can u pls explain how u are able to train with 5 fold CV within 9 hrs on kaggle using full data ? </p>",
          "rawMarkdown": "@drhabib  Hi. Are u training locally ?  If no, then can u pls explain how u are able to train with 5 fold CV within 9 hrs on kaggle using full data ? ",
          "votes": 3
        },
        {
          "id": 705828,
          "postDate": "2019-12-29T14:48:10.173Z",
          "content": "<p>Yes I am training locally :) BTW per competition rules you are allowed to use only 2h of gpu and 9h of cpu when inferring. Currently my inference with 5 fold cv takes 1h2min (still have to optimise)</p>",
          "rawMarkdown": "Yes I am training locally :) BTW per competition rules you are allowed to use only 2h of gpu and 9h of cpu when inferring. Currently my inference with 5 fold cv takes 1h2min (still have to optimise)",
          "votes": 4
        },
        {
          "id": 738999,
          "postDate": "2020-02-07T09:15:37.720Z",
          "content": "<p>Locally :). What is your machine's configuration? RAM/GPU/Cost/Vendor</p>",
          "rawMarkdown": "Locally :). What is your machine's configuration? RAM/GPU/Cost/Vendor"
        }
      ]
    },
    {
      "id": 704990,
      "postDate": "2019-12-28T08:59:25.567Z",
      "content": "<p>I published <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">Bengali: SEResNeXt prediction with pytorch</a> which achieves 0.9583.</p>\n\n<p>```\nmodel: seresnext101_32x4d (pretrained=True)\nimg: 1 channel \nimg_sz: 128\nsplit: random (90/10)\noptim: Adam</p>\n\n<p>LB:  0.9583\n```</p>\n\n<p>This is my first submit and since I have not tuned so much, I think the score can be improved more.</p>",
      "rawMarkdown": "I published [Bengali: SEResNeXt prediction with pytorch](https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch) which achieves 0.9583.\n\n```\nmodel: seresnext101_32x4d (pretrained=True)\nimg: 1 channel \nimg_sz: 128\nsplit: random (90/10)\noptim: Adam\n\nLB:  0.9583\n```\n\nThis is my first submit and since I have not tuned so much, I think the score can be improved more.",
      "votes": 3,
      "replies": [
        {
          "id": 706105,
          "postDate": "2019-12-29T23:27:43.887Z",
          "content": "<p>Kernel updated.</p>\n\n<p>```\nmodel: seresnext50_32x4d (pretrained=True)\nimg: 1 channel \nimg_sz: 128\nsplit: random (90/10)\noptim: Adam</p>\n\n<p>LB:  0.9607\n```</p>\n\n<p>(since I modified training/prediction script a bit, its difference is not only the model.)</p>",
          "rawMarkdown": "Kernel updated.\n\n```\nmodel: seresnext50_32x4d (pretrained=True)\nimg: 1 channel \nimg_sz: 128\nsplit: random (90/10)\noptim: Adam\n\nLB:  0.9607\n```\n\n(since I modified training/prediction script a bit, its difference is not only the model.)",
          "votes": 2
        }
      ]
    },
    {
      "id": 703421,
      "postDate": "2019-12-26T06:16:57.530Z",
      "content": "<p>model: resnet34(pretrained=False)\nimg: 1 channel \nimg_sz: 112\nsplit: random (50/50)- trained only on 168*5 images per class\noptim: Adam\nepoch: 100 +EarlyStopping ( patience = 10)\nsched: -\npreprocesssing - cropping to character + resize to 112x112\nCV : 0.9588\nLB:  0.9373</p>",
      "rawMarkdown": "model: resnet34(pretrained=False)\nimg: 1 channel \nimg_sz: 112\nsplit: random (50/50)- trained only on 168*5 images per class\noptim: Adam\nepoch: 100 +EarlyStopping ( patience = 10)\nsched: -\npreprocesssing - cropping to character + resize to 112x112\nCV : 0.9588\nLB:  0.9373",
      "votes": 3
    },
    {
      "id": 710751,
      "postDate": "2020-01-05T07:03:47.857Z",
      "content": "<p>I got 0.9671 on public LB with a single DenseNet121.</p>\n\n<p><code>\nmodel: DenseNet-121 (modified some parts)\nimg: 1 channel \ndim: 128\nsplit: random 80/20\noptim: SGD\nepoch: 60\n</code></p>",
      "rawMarkdown": "I got 0.9671 on public LB with a single DenseNet121.\n\n```\nmodel: DenseNet-121 (modified some parts)\nimg: 1 channel \ndim: 128\nsplit: random 80/20\noptim: SGD\nepoch: 60\n```",
      "votes": 4,
      "replies": [
        {
          "id": 710758,
          "postDate": "2020-01-05T07:17:10.543Z",
          "content": "<p>Great! What was your CV score?</p>",
          "rawMarkdown": "Great! What was your CV score?",
          "votes": 1
        }
      ]
    },
    {
      "id": 707575,
      "postDate": "2020-01-01T04:58:20.913Z",
      "content": "<p>model: resnet34(pretrained=True)\nimg: 3 channel \nimg_sz: 128</p>\n\n<p>CV : 0.9650\nLB:  0.9670</p>",
      "rawMarkdown": "model: resnet34(pretrained=True)\nimg: 3 channel \nimg_sz: 128\n\nCV : 0.9650\nLB:  0.9670",
      "votes": 4,
      "replies": [
        {
          "id": 707794,
          "postDate": "2020-01-01T13:58:52.457Z",
          "content": "<p><a href=\"/garybios\">@garybios</a>  Hi. I have few questions:\n1. Img = 3 channel.. are u just copying the 1 channel 3 times? \n2. Is it ur 5 fold cv? If yes, for how many epochs are u training each fold ? \n3. What is ur learning rate and scheduler ? </p>\n\n<p>Thanks </p>",
          "rawMarkdown": "@garybios  Hi. I have few questions:\n1. Img = 3 channel.. are u just copying the 1 channel 3 times? \n2. Is it ur 5 fold cv? If yes, for how many epochs are u training each fold ? \n3. What is ur learning rate and scheduler ? \n\nThanks ",
          "votes": 2
        }
      ]
    },
    {
      "id": 704943,
      "postDate": "2019-12-28T07:30:34.463Z",
      "content": "<p>Hey Folks, \nAm i missing something or data seems to be \"easy\" that everyone has so high metric scores straightaway? \nPlus LB in somewhat sync with CV stealing the fun!\nPS Newbie to this comp..</p>",
      "rawMarkdown": "Hey Folks, \nAm i missing something or data seems to be \"easy\" that everyone has so high metric scores straightaway? \nPlus LB in somewhat sync with CV stealing the fun!\nPS Newbie to this comp..",
      "votes": 4,
      "replies": [
        {
          "id": 707207,
          "postDate": "2019-12-31T12:17:24.137Z",
          "content": "<p>Since the competition is hosted by an AI community ,I think they know the nuisances and taken some extra care to give reasonably good data . I think there are no easy comp when you are putting effort to try and get a medal by competing with your Kaggle gurus . There are also lot of opportunities to try out new papers in this comp . </p>",
          "rawMarkdown": "Since the competition is hosted by an AI community ,I think they know the nuisances and taken some extra care to give reasonably good data . I think there are no easy comp when you are putting effort to try and get a medal by competing with your Kaggle gurus . There are also lot of opportunities to try out new papers in this comp . ",
          "votes": 3
        }
      ]
    },
    {
      "id": 703276,
      "postDate": "2019-12-25T23:12:22.040Z",
      "content": "<p>How did you make the model accept gray scale images? From keras.applications, they only accept RGB images.</p>",
      "rawMarkdown": "How did you make the model accept gray scale images? From keras.applications, they only accept RGB images.",
      "votes": 4,
      "replies": [
        {
          "id": 703414,
          "postDate": "2019-12-26T06:12:51.777Z",
          "content": "<p>In the very first nn.Conv2d in the first residual block , change the first parameter (in_channels) to 1 I guess? That's how we did it , however we wrote the complete ResNet architecture rather than import it so that it gives you fine tuning options like these.</p>",
          "rawMarkdown": "In the very first nn.Conv2d in the first residual block , change the first parameter (in_channels) to 1 I guess? That's how we did it , however we wrote the complete ResNet architecture rather than import it so that it gives you fine tuning options like these.",
          "votes": 1
        },
        {
          "id": 703420,
          "postDate": "2019-12-26T06:16:02.540Z",
          "content": "<p>Cool! I saw some starter kernel added a conv2d block with out_channels 3 and in_channels 1. </p>",
          "rawMarkdown": "Cool! I saw some starter kernel added a conv2d block with out_channels 3 and in_channels 1. ",
          "votes": 1
        },
        {
          "id": 703423,
          "postDate": "2019-12-26T06:18:41.503Z",
          "content": "<p>Yeah you can just add a Conv2d to like that then pass it to the neural net too!</p>",
          "rawMarkdown": "Yeah you can just add a Conv2d to like that then pass it to the neural net too!",
          "votes": 1
        },
        {
          "id": 703425,
          "postDate": "2019-12-26T06:20:38.160Z",
          "content": "<p>I will try it soon. I am still waiting for my original xception network to train and see what we got here for LB.</p>",
          "rawMarkdown": "I will try it soon. I am still waiting for my original xception network to train and see what we got here for LB.",
          "votes": 1
        },
        {
          "id": 705460,
          "postDate": "2019-12-29T00:09:26.290Z",
          "content": "<p>in Pytorch you can do this small trick , convert any pretrain network to accept 1 channel images without loosing pertained weights.</p>\n\n<p><code>\narch = models.resnet50(num_classes=1000, pretrained=True)\narch = list(arch.children())\nw = arch[0].weight\narch[0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=2, bias=False)\narch[0].weight = nn.Parameter(torch.mean(w, dim=1, keepdim=True))\narch = nn.Sequential(*arch)\n</code></p>\n\n<p>Bassicly we are taking weights of the first <code>Conv2d</code> storing them in <code>w</code>. Afterwards we are creating new <code>Conv2d</code> which has 1 channel and replacing its parameter with <code>w</code> for which we take mean. I am pretty sure this can be achieved in <code>keras</code> or <code>tenserflow</code></p>",
          "rawMarkdown": "in Pytorch you can do this small trick , convert any pretrain network to accept 1 channel images without loosing pertained weights.\n\n```\narch = models.resnet50(num_classes=1000, pretrained=True)\narch = list(arch.children())\nw = arch[0].weight\narch[0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=2, bias=False)\narch[0].weight = nn.Parameter(torch.mean(w, dim=1, keepdim=True))\narch = nn.Sequential(*arch)\n```\n\nBassicly we are taking weights of the first `Conv2d` storing them in `w`. Afterwards we are creating new `Conv2d` which has 1 channel and replacing its parameter with `w` for which we take mean. I am pretty sure this can be achieved in `keras` or `tenserflow`",
          "votes": 19
        }
      ]
    },
    {
      "id": 759655,
      "postDate": "2020-02-29T09:35:24.437Z",
      "content": "<p>(update)Model: Seresnet50\nimg_size: 3x137 * 236\nAugmentation:  cutmix + cutout\nCV: 0.997\nLB: 0.989</p>",
      "rawMarkdown": "(update)Model: Seresnet50\nimg_size: 3x137 * 236\nAugmentation:  cutmix + cutout\nCV: 0.997\nLB: 0.989",
      "votes": 3,
      "replies": [
        {
          "id": 759898,
          "postDate": "2020-02-29T15:40:39.457Z",
          "content": "<p>Great! How is it possible to train a model for 110 epochs? How does not it overfit?</p>",
          "rawMarkdown": "Great! How is it possible to train a model for 110 epochs? How does not it overfit?"
        },
        {
          "id": 759976,
          "postDate": "2020-02-29T17:21:41.200Z",
          "content": "<p>\" How does not it overfit?\"</p>\n\n<p>augmentation. </p>",
          "rawMarkdown": "\" How does not it overfit?\"\n\naugmentation. ",
          "votes": 3
        },
        {
          "id": 763042,
          "postDate": "2020-03-04T04:13:54.303Z",
          "content": "<p>Curious why did you decide to go for 128x128 and not 137x236? :)</p>",
          "rawMarkdown": "Curious why did you decide to go for 128x128 and not 137x236? :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 741811,
      "postDate": "2020-02-11T00:42:50.960Z",
      "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a> I was wondering if you could share your code for cv 0.9630 lb 0.9611 model. I wonder how you could make the gap between cv and lb so small. I currently do 80/20 stratified shuffle split and reached 0.9886 cv, but lb is only 0.9679, so I would really appreciate if you could publish you low-score pipeline. Thank you very much!</p>",
      "rawMarkdown": "Hi @drhabib I was wondering if you could share your code for cv 0.9630 lb 0.9611 model. I wonder how you could make the gap between cv and lb so small. I currently do 80/20 stratified shuffle split and reached 0.9886 cv, but lb is only 0.9679, so I would really appreciate if you could publish you low-score pipeline. Thank you very much!",
      "votes": 2
    },
    {
      "id": 768182,
      "postDate": "2020-03-10T14:17:22.283Z",
      "content": "<p>model: SE-ResNeXt50 with DropBlock\ninput size: 128x128\nvalidation: iterative stratification (single fold)</p>\n\n<p>CV=0.9888, LB=0.9758\nCV=0.9905, LB=0.9764</p>\n\n<p>Large CV/LB gap...</p>",
      "rawMarkdown": "model: SE-ResNeXt50 with DropBlock\ninput size: 128x128\nvalidation: iterative stratification (single fold)\n\nCV=0.9888, LB=0.9758\nCV=0.9905, LB=0.9764\n\nLarge CV/LB gap...",
      "votes": 1,
      "replies": [
        {
          "id": 768247,
          "postDate": "2020-03-10T15:26:01.873Z",
          "content": "<p>Hey im also using dropblock as augmentations, to reduce the gap i found that having less dropblock layers and adding other augmentations like cutout and rotate gave me a better score! My gap is a lot smaller since then, cv 0.9883 lb 0.9825</p>",
          "rawMarkdown": "Hey im also using dropblock as augmentations, to reduce the gap i found that having less dropblock layers and adding other augmentations like cutout and rotate gave me a better score! My gap is a lot smaller since then, cv 0.9883 lb 0.9825"
        },
        {
          "id": 768411,
          "postDate": "2020-03-10T19:04:23.537Z",
          "content": "<p>Hi! Probably you are using wrong metrics how is mentioned here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/133219\">https://www.kaggle.com/c/bengaliai-cv19/discussion/133219</a> . It was the reason for my big gap )</p>",
          "rawMarkdown": "Hi! Probably you are using wrong metrics how is mentioned here: https://www.kaggle.com/c/bengaliai-cv19/discussion/133219 . It was the reason for my big gap )"
        },
        {
          "id": 769188,
          "postDate": "2020-03-11T16:12:34.560Z",
          "content": "<p>I ask a silly question,what's the meaning of  iterative stratification and how to implement it?</p>",
          "rawMarkdown": "I ask a silly question,what's the meaning of  iterative stratification and how to implement it?",
          "votes": -1
        },
        {
          "id": 769235,
          "postDate": "2020-03-11T17:06:36Z",
          "content": "<p><a href=\"/thefatcat\">@thefatcat</a> these <a href=\"https://www.kaggle.com/yiheng/iterative-stratification\">https://www.kaggle.com/yiheng/iterative-stratification</a> <a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a> might help.</p>",
          "rawMarkdown": "@thefatcat these https://www.kaggle.com/yiheng/iterative-stratification https://github.com/trent-b/iterative-stratification might help."
        }
      ]
    },
    {
      "id": 761813,
      "postDate": "2020-03-03T00:29:54.730Z",
      "content": "<p>Model: DenseNet121\nSize: 224x224\nAug: CO, GM\nCV: 0.9936\nLB: 0.9842</p>\n\n<p>I still have few ideas that I'm experimenting with right now.\nI might be looking for a teammate with a high scoring seresnext50 model (I didn't have any luck with it).\nPM me if anyone would be interested</p>",
      "rawMarkdown": "Model: DenseNet121\nSize: 224x224\nAug: CO, GM\nCV: 0.9936\nLB: 0.9842\n\nI still have few ideas that I'm experimenting with right now.\nI might be looking for a teammate with a high scoring seresnext50 model (I didn't have any luck with it).\nPM me if anyone would be interested",
      "votes": 1,
      "replies": [
        {
          "id": 768222,
          "postDate": "2020-03-10T14:58:19.763Z",
          "content": "<p>What is GM?</p>",
          "rawMarkdown": "What is GM?"
        },
        {
          "id": 768330,
          "postDate": "2020-03-10T16:42:55.663Z",
          "content": "<p><a href=\"/gxygomes\">@gxygomes</a> I assume it is GridMask</p>",
          "rawMarkdown": "@gxygomes I assume it is GridMask"
        }
      ]
    },
    {
      "id": 761792,
      "postDate": "2020-03-03T00:00:39.750Z",
      "content": "<p>Model: se-resnext50\nimgSize: 128x128\nAug: Cutout\nEpoch: 200\nLocal validation 0.9921\nLB: 0.9815</p>",
      "rawMarkdown": "Model: se-resnext50\nimgSize: 128x128\nAug: Cutout\nEpoch: 200\nLocal validation 0.9921\nLB: 0.9815",
      "votes": 1,
      "replies": [
        {
          "id": 764209,
          "postDate": "2020-03-05T08:17:21.170Z",
          "content": "<p>May I ask the minimum lr you reach</p>",
          "rawMarkdown": "May I ask the minimum lr you reach"
        },
        {
          "id": 764227,
          "postDate": "2020-03-05T08:28:52.140Z",
          "content": "<p>1e-6</p>",
          "rawMarkdown": "1e-6"
        },
        {
          "id": 768333,
          "postDate": "2020-03-10T16:46:44.437Z",
          "content": "<p><a href=\"/anthonymarcz\">@anthonymarcz</a> can i know is cutout was your only augmentation?? Are you using multiple cuts? And what was your height and width of cuts?</p>",
          "rawMarkdown": "@anthonymarcz can i know is cutout was your only augmentation?? Are you using multiple cuts? And what was your height and width of cuts?"
        }
      ]
    },
    {
      "id": 725802,
      "postDate": "2020-01-22T14:06:59.213Z",
      "content": "<p>CV: 97.03%\nLB: 95.64%</p>\n\n<p><code>\nmodel: DeepCNN\nimg: 1 channel \nimg_sz: 64\nsplit: random (80/20)\noptim: Adam\nepoch: 15\nrotate:20\n</code></p>",
      "rawMarkdown": "CV: 97.03%\nLB: 95.64%\n\n```\nmodel: DeepCNN\nimg: 1 channel \nimg_sz: 64\nsplit: random (80/20)\noptim: Adam\nepoch: 15\nrotate:20\n```\n",
      "votes": 1
    },
    {
      "id": 722046,
      "postDate": "2020-01-18T03:11:52.317Z",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> Hi! Thank you for sharing your wonderful results. Just wondering, when doing cutmix / mixup is there probability of using usual training at all? So far as I know, implementation of mixup in fast.ai only uses mixup with .4 probability. With \nif np.random.rand()&lt;0.5: do mixup\nelse: do cutmix\nThe model is purely trained using samples mixed with mixup or cutmix?</p>",
      "rawMarkdown": "@bibek777 Hi! Thank you for sharing your wonderful results. Just wondering, when doing cutmix / mixup is there probability of using usual training at all? So far as I know, implementation of mixup in fast.ai only uses mixup with .4 probability. With \nif np.random.rand()&lt;0.5: do mixup\nelse: do cutmix\nThe model is purely trained using samples mixed with mixup or cutmix?",
      "votes": 1
    },
    {
      "id": 711664,
      "postDate": "2020-01-06T11:19:07.270Z",
      "content": "<p>eff0 \ncv5: 0.9689\nlb:  0.9615</p>",
      "rawMarkdown": "eff0 \ncv5: 0.9689\nlb:  0.9615",
      "votes": 1,
      "replies": [
        {
          "id": 717443,
          "postDate": "2020-01-13T05:43:43.117Z",
          "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a>, Nice result! How many epochs did you run? For me, I can only get to LB: 0.954 after 10-12 epochs(full model training). </p>",
          "rawMarkdown": "@cswwp347724, Nice result! How many epochs did you run? For me, I can only get to LB: 0.954 after 10-12 epochs(full model training). "
        },
        {
          "id": 717454,
          "postDate": "2020-01-13T06:15:23.807Z",
          "content": "<p>It's cv5 result with eff0, my single eff0 best result is 0.9592 with few train data modification</p>",
          "rawMarkdown": "It's cv5 result with eff0, my single eff0 best result is 0.9592 with few train data modification"
        }
      ]
    },
    {
      "id": 708473,
      "postDate": "2020-01-02T10:29:49.577Z",
      "content": "<p>one of my tries</p>\n\n<p><code>\nmodel: B0(pretrained=True)\nimg: 3 channel\nimg_sz: 256\nsplit: random split (80/20)\noptim: Over9000\nepoch: 20\nsched: Cosine w/o warmup\n</code></p>\n\n<p>I also tried B2 and B4 models but they didn't get a boost. \nMaybe more tuning is needed or just model is just big enough to overfit.</p>\n\n<p>CV : 0.9703\nLB : 0.9640</p>",
      "rawMarkdown": "one of my tries\n\n```\nmodel: B0(pretrained=True)\nimg: 3 channel\nimg_sz: 256\nsplit: random split (80/20)\noptim: Over9000\nepoch: 20\nsched: Cosine w/o warmup\n```\n\nI also tried B2 and B4 models but they didn't get a boost. \nMaybe more tuning is needed or just model is just big enough to overfit.\n\nCV : 0.9703\nLB : 0.9640",
      "votes": 1,
      "replies": [
        {
          "id": 708487,
          "postDate": "2020-01-02T10:50:28.597Z",
          "content": "<p>Thanks for sharing, are you using pytorch or tensorflow?</p>",
          "rawMarkdown": "Thanks for sharing, are you using pytorch or tensorflow?",
          "votes": 1
        },
        {
          "id": 708574,
          "postDate": "2020-01-02T12:51:21.160Z",
          "content": "<p>i use pytorch!</p>",
          "rawMarkdown": "i use pytorch!",
          "votes": 1
        },
        {
          "id": 709230,
          "postDate": "2020-01-03T07:57:13.150Z",
          "content": "<p>Thanks, may I ask which implementation you are using?\nI've been trying B0 with this implementation <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch\">https://github.com/lukemelas/EfficientNet-PyTorch</a>.\nBut with the same stack one epoch takes about 9min while one epoch for ResNet18 takes about 3min, I find this a bit strange since B0 is supposed to be ultra light. Have you noticed something similar?\nI twicked the code a little to get 3 heads like this one afet the extract_features part : \n```\nclass Head(torch.nn.Module):\n    def <strong>init</strong>(self, out_features, out_channels, dropout=0.2):\n        super(Head, self).<strong>init</strong>()\n        self._avg_pooling = torch.nn.AdaptiveAvgPool2d(1)\n        self._dropout = torch.nn.Dropout(dropout)\n        self._fc = torch.nn.Linear(out_channels, out_features)</p>\n\n<pre><code>def forward(self, x, batch_size):\n    x = self._avg_pooling(x)\n    x = x.view(batch_size, -1)\n    x = self._dropout(x)\n    x = self._fc(x)\n    return x\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "Thanks, may I ask which implementation you are using?\nI've been trying B0 with this implementation https://github.com/lukemelas/EfficientNet-PyTorch.\nBut with the same stack one epoch takes about 9min while one epoch for ResNet18 takes about 3min, I find this a bit strange since B0 is supposed to be ultra light. Have you noticed something similar?\nI twicked the code a little to get 3 heads like this one afet the extract_features part : \n```\nclass Head(torch.nn.Module):\n    def __init__(self, out_features, out_channels, dropout=0.2):\n        super(Head, self).__init__()\n        self._avg_pooling = torch.nn.AdaptiveAvgPool2d(1)\n        self._dropout = torch.nn.Dropout(dropout)\n        self._fc = torch.nn.Linear(out_channels, out_features)\n        \n    def forward(self, x, batch_size):\n        x = self._avg_pooling(x)\n        x = x.view(batch_size, -1)\n        x = self._dropout(x)\n        x = self._fc(x)\n        return x\n```",
          "votes": 1
        },
        {
          "id": 709267,
          "postDate": "2020-01-03T09:50:13.530Z",
          "content": "<p>my implementation is based on <code>efficientnet_pytorch</code> repo, maybe similar w/ referred notebook.\nand i'm using colab (P100) for training, i just took <code>20min</code> for one epoch, B0 and <code>16min</code> for one epoch, ResNet34.</p>\n\n<p>maybe there're many reasons for slow training time, cuz of data pipelining, gpu utilization, model size &amp; flops, callbacks etc...\nand also training time is not directly proportional w/ not only the number of model parameters!</p>\n\n<p>maybe still matrix multiplication (e.g. Head part) could be the one of the reason.\n<code>B0</code>'s bottleneck feature size is <code>1280</code> and <code>ResNet18</code> is <code>512</code>.</p>\n\n<p>consider all of above's artifacts, maybe the training time could be reasonable!</p>",
          "rawMarkdown": "my implementation is based on `efficientnet_pytorch` repo, maybe similar w/ referred notebook.\nand i'm using colab (P100) for training, i just took `20min` for one epoch, B0 and `16min` for one epoch, ResNet34.\n\nmaybe there're many reasons for slow training time, cuz of data pipelining, gpu utilization, model size &amp; flops, callbacks etc...\nand also training time is not directly proportional w/ not only the number of model parameters!\n\nmaybe still matrix multiplication (e.g. Head part) could be the one of the reason.\n`B0`'s bottleneck feature size is `1280` and `ResNet18` is `512`.\n\nconsider all of above's artifacts, maybe the training time could be reasonable!",
          "votes": 3
        },
        {
          "id": 709312,
          "postDate": "2020-01-03T10:52:51.317Z",
          "content": "<p>thanks, so it seems that 9min per epoch is reasonnable! I tried resnet34 and training one epoch (on 80% of the training data) takes ~5min for me, so B0 is indeed even slower than resnet34</p>",
          "rawMarkdown": "thanks, so it seems that 9min per epoch is reasonnable! I tried resnet34 and training one epoch (on 80% of the training data) takes ~5min for me, so B0 is indeed even slower than resnet34",
          "votes": 1
        }
      ]
    },
    {
      "id": 770620,
      "postDate": "2020-03-13T06:58:28.727Z",
      "content": "<p>augmentation: Cutmix + cutout + grid\ninput size: 137 x 236\nsplit: 97.5% + 2.5%</p>\n\n<p>Model/CV/LB: SE-ResNext50/99.38/98.57 with OHEM\nModel/CV/LB: SE-ResNext50/99.59/98.54 w/o OHEM\nModel/CV/LB: B3/99.20/98.18 with OHEM</p>",
      "rawMarkdown": "augmentation: Cutmix + cutout + grid\ninput size: 137 x 236\nsplit: 97.5% + 2.5%\n\nModel/CV/LB: SE-ResNext50/99.38/98.57 with OHEM\nModel/CV/LB: SE-ResNext50/99.59/98.54 w/o OHEM\nModel/CV/LB: B3/99.20/98.18 with OHEM",
      "votes": 2
    },
    {
      "id": 768540,
      "postDate": "2020-03-10T23:54:07.830Z",
      "content": "<p>i wondered if anyone has results with models using dilated convolution?</p>",
      "rawMarkdown": "i wondered if anyone has results with models using dilated convolution?",
      "votes": 2
    },
    {
      "id": 763422,
      "postDate": "2020-03-04T13:22:45.220Z",
      "content": "<p>resnet34: single fold(all train data)\nLB: 0.9807</p>\n\n<p>EfficientNet-b2: iterative stratification(train:valid=8:2), not 5 fold average\nCV: 0.9916\nLB: 0.9865</p>\n\n<p>I'm afraid of overfitting. Have a nice day!</p>",
      "rawMarkdown": "resnet34: single fold(all train data)\nLB: 0.9807\n\nEfficientNet-b2: iterative stratification(train:valid=8:2), not 5 fold average\nCV: 0.9916\nLB: 0.9865\n\nI'm afraid of overfitting. Have a nice day!",
      "votes": 2,
      "replies": [
        {
          "id": 763427,
          "postDate": "2020-03-04T13:33:20.733Z",
          "content": "<p><a href=\"/kyosato\">@kyosato</a> May I ask a small question, is it just changing the network then get a big improvement ?</p>",
          "rawMarkdown": "@kyosato May I ask a small question, is it just changing the network then get a big improvement ?",
          "votes": 2
        },
        {
          "id": 763460,
          "postDate": "2020-03-04T13:59:38.673Z",
          "content": "<p><a href=\"/hesene\">@hesene</a> \nAlmost Yes, except for fold strategy and epoch number.</p>",
          "rawMarkdown": "@hesene \nAlmost Yes, except for fold strategy and epoch number.",
          "votes": 1
        },
        {
          "id": 763466,
          "postDate": "2020-03-04T14:05:37.740Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks",
          "votes": 1
        },
        {
          "id": 763783,
          "postDate": "2020-03-04T21:08:08.787Z",
          "content": "<p><a href=\"/kyosato\">@kyosato</a> how many epochs did you train for?</p>",
          "rawMarkdown": "@kyosato how many epochs did you train for?"
        },
        {
          "id": 763798,
          "postDate": "2020-03-04T21:51:02.837Z",
          "content": "<p>over 100</p>",
          "rawMarkdown": "over 100"
        },
        {
          "id": 764396,
          "postDate": "2020-03-05T12:21:09.853Z",
          "content": "<p>Astrologers have announced a week of experiments with EfficientNet-b2 😃 </p>",
          "rawMarkdown": "Astrologers have announced a week of experiments with EfficientNet-b2 😃 "
        },
        {
          "id": 764483,
          "postDate": "2020-03-05T14:05:00.573Z",
          "content": "<p><a href=\"/kyosato\">@kyosato</a>  what augmentations you tried?</p>",
          "rawMarkdown": "@kyosato  what augmentations you tried?"
        },
        {
          "id": 764552,
          "postDate": "2020-03-05T15:34:26.183Z",
          "content": "<p>How you train model over 100 epoch without overfitting? what lr do you use? Adam or SGD?</p>",
          "rawMarkdown": "How you train model over 100 epoch without overfitting? what lr do you use? Adam or SGD?",
          "votes": 1
        },
        {
          "id": 764808,
          "postDate": "2020-03-05T23:31:57.757Z",
          "content": "<p>Sorry, I can't tell you the detail, but Cutmix is very good.</p>",
          "rawMarkdown": "Sorry, I can't tell you the detail, but Cutmix is very good.",
          "votes": 1
        }
      ]
    },
    {
      "id": 761109,
      "postDate": "2020-03-02T06:37:36.253Z",
      "content": "<p>model:seresnext50\nsize:3x137x236\naugumentation: cutmix etc\nCV:0.9965\nLB:0.9866\nLooking for a teammate!</p>",
      "rawMarkdown": "model:seresnext50\nsize:3x137x236\naugumentation: cutmix etc\nCV:0.9965\nLB:0.9866\nLooking for a teammate!",
      "votes": 2,
      "replies": [
        {
          "id": 761114,
          "postDate": "2020-03-02T06:48:37.687Z",
          "content": "<p>That is a really amazing result! Have you tried even larger model or larger image size (224*224)?</p>",
          "rawMarkdown": "That is a really amazing result! Have you tried even larger model or larger image size (224*224)?"
        },
        {
          "id": 761127,
          "postDate": "2020-03-02T07:07:05.600Z",
          "content": "<p>Thank you for your reply.\nI tried the size of 224x224 once, but it did not work.\nI am thinking about why it happened.</p>",
          "rawMarkdown": "Thank you for your reply.\nI tried the size of 224x224 once, but it did not work.\nI am thinking about why it happened.",
          "votes": 3
        },
        {
          "id": 761845,
          "postDate": "2020-03-03T01:33:00.387Z",
          "content": "<p>Me too, spent two days going nowhere with it ;(</p>",
          "rawMarkdown": "Me too, spent two days going nowhere with it ;(",
          "votes": 2
        }
      ]
    },
    {
      "id": 757530,
      "postDate": "2020-02-26T21:53:16.410Z",
      "content": "<p>Model: EfficientNetB3\nimg_size: 3x102x177 (just scaled with factor 0.75)\nAugmentation: No Augmentation\nEpoch: 60\nCV: 0.977 (using validation score for root recall..I noticed this gives a way better indication then the validation score for the official metric.\nLB: 0.9724</p>",
      "rawMarkdown": "Model: EfficientNetB3\nimg_size: 3x102x177 (just scaled with factor 0.75)\nAugmentation: No Augmentation\nEpoch: 60\nCV: 0.977 (using validation score for root recall..I noticed this gives a way better indication then the validation score for the official metric.\nLB: 0.9724",
      "votes": 2,
      "replies": [
        {
          "id": 759003,
          "postDate": "2020-02-28T12:48:42.113Z",
          "content": "<p><a href=\"/rsmits\">@rsmits</a> Very charming results using no augmentation! Really getting me curious how you can do it without augmentation, tho</p>",
          "rawMarkdown": "@rsmits Very charming results using no augmentation! Really getting me curious how you can do it without augmentation, tho",
          "votes": 1
        },
        {
          "id": 759016,
          "postDate": "2020-02-28T13:03:43.307Z",
          "content": "<p>Hi <a href=\"/roguekk007\">@roguekk007</a> Thank you for the compliment. How I do it...well that's for the most part no secret ;-)\nCheck out <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974\">https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974</a> the discussion with Heng CherKeng.</p>\n\n<p>The method I used for training is in my github. That gave my public kernel a score of 0.9703. An updated version of that code already pushed me to 0.9728. \nI think I can get even higher...but I have to wait for my computer to finish the epochs ;-)</p>",
          "rawMarkdown": "Hi @roguekk007 Thank you for the compliment. How I do it...well that's for the most part no secret ;-)\nCheck out [https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974](https://www.kaggle.com/c/bengaliai-cv19/discussion/132228#758974) the discussion with Heng CherKeng.\n\nThe method I used for training is in my github. That gave my public kernel a score of 0.9703. An updated version of that code already pushed me to 0.9728. \nI think I can get even higher...but I have to wait for my computer to finish the epochs ;-)"
        },
        {
          "id": 763291,
          "postDate": "2020-03-04T10:27:38.760Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 724404,
      "postDate": "2020-01-21T06:07:49.580Z",
      "content": "<p>i wonder is there any comparison with and without using imagenet pretrained models?\ni have been using pretrained models for my work and is curious about the performance for training from scratch</p>",
      "rawMarkdown": "i wonder is there any comparison with and without using imagenet pretrained models?\ni have been using pretrained models for my work and is curious about the performance for training from scratch",
      "votes": 2,
      "replies": [
        {
          "id": 724444,
          "postDate": "2020-01-21T06:45:55.470Z",
          "content": "<p>We did use ResNet-34 and 50 ,earlier trained from scratch for ~10 epochs , but pre-trained performed better for the same number of epochs. Maybe training from scratch for ~40+ epochs might work better than pre-trained?</p>",
          "rawMarkdown": "We did use ResNet-34 and 50 ,earlier trained from scratch for ~10 epochs , but pre-trained performed better for the same number of epochs. Maybe training from scratch for ~40+ epochs might work better than pre-trained?",
          "votes": 2
        },
        {
          "id": 730314,
          "postDate": "2020-01-27T10:42:39.937Z",
          "content": "<p>For me, some of layers use pretrained weights and it works best. If I load all pretrained weights for all layers, it degrades the performance a bit.</p>",
          "rawMarkdown": "For me, some of layers use pretrained weights and it works best. If I load all pretrained weights for all layers, it degrades the performance a bit.",
          "votes": 8
        }
      ]
    },
    {
      "id": 710682,
      "postDate": "2020-01-05T04:38:50.803Z",
      "content": "<p><code>\nmodel: resnet18\nimg: 1 channel \ndim: 128\nsplit: iterative stratified split(80/20)\noptim: over9000\nepoch: 32\nCV : 0.9676\nLB:  0.9604\n</code>\nThe gap between my CV and LB is always around 0.007. Any suggestion regarding train-val split to get similar scores would be appreciated. </p>",
      "rawMarkdown": "```\nmodel: resnet18\nimg: 1 channel \ndim: 128\nsplit: iterative stratified split(80/20)\noptim: over9000\nepoch: 32\nCV : 0.9676\nLB:  0.9604\n```\nThe gap between my CV and LB is always around 0.007. Any suggestion regarding train-val split to get similar scores would be appreciated. ",
      "votes": 2
    },
    {
      "id": 1006617,
      "postDate": "2020-09-11T12:21:21.710Z",
      "content": "<p>nb…………</p>",
      "rawMarkdown": "nb............"
    },
    {
      "id": 767900,
      "postDate": "2020-03-10T08:17:02.373Z",
      "content": "<p>CV 0.988 LB ?\nI'm new to this competition, still cannot submit Lol</p>",
      "rawMarkdown": "CV 0.988 LB ?\nI'm new to this competition, still cannot submit Lol",
      "replies": [
        {
          "id": 768211,
          "postDate": "2020-03-10T14:46:52.207Z",
          "content": "<p>For me, CV0.9872 gave LB0.9793. I think you will get around LB0.98 :)</p>",
          "rawMarkdown": "For me, CV0.9872 gave LB0.9793. I think you will get around LB0.98 :)",
          "votes": 2
        },
        {
          "id": 769753,
          "postDate": "2020-03-12T08:19:55.817Z",
          "content": "<p>Thanks !! ;)\nI'll try to submit, maybe tomorrow😄 </p>",
          "rawMarkdown": "Thanks !! ;)\nI'll try to submit, maybe tomorrow😄 "
        },
        {
          "id": 769756,
          "postDate": "2020-03-12T08:20:54.423Z",
          "content": "<p>Update : CV 0.992 LB ?\nI cannot waste sub then LB is still misterious....</p>",
          "rawMarkdown": "Update : CV 0.992 LB ?\nI cannot waste sub then LB is still misterious...."
        },
        {
          "id": 771247,
          "postDate": "2020-03-13T23:27:19.523Z",
          "content": "<p>Update : CV 0.993 LB ?</p>",
          "rawMarkdown": "Update : CV 0.993 LB ?",
          "votes": -2
        }
      ]
    },
    {
      "id": 762046,
      "postDate": "2020-03-03T06:44:29.980Z",
      "content": "<p>CV : 0.9775\nLB:  0.9747\nseresnext50\nStill trying to use different augs, looking forward to improve score more</p>",
      "rawMarkdown": "CV : 0.9775\nLB:  0.9747\nseresnext50\nStill trying to use different augs, looking forward to improve score more"
    },
    {
      "id": 753070,
      "postDate": "2020-02-21T17:19:43.697Z",
      "content": "<p>Is there any notebook that uses cutmix in PyTorch. There are no such notebooks in the kernel section. Can anyone share?</p>",
      "rawMarkdown": "Is there any notebook that uses cutmix in PyTorch. There are no such notebooks in the kernel section. Can anyone share?",
      "replies": [
        {
          "id": 753127,
          "postDate": "2020-02-21T18:38:03.577Z",
          "content": "<p>Look in discussions</p>",
          "rawMarkdown": "Look in discussions"
        }
      ]
    },
    {
      "id": 761066,
      "postDate": "2020-03-02T04:59:15.577Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 759957,
      "postDate": "2020-02-29T16:49:44.367Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 742605,
      "postDate": "2020-02-11T12:07:23.333Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 704188,
      "postDate": "2019-12-27T06:32:19.100Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 710146,
      "postDate": "2020-01-04T11:05:38.943Z",
      "content": "<p>Thanks for sharing... </p>",
      "rawMarkdown": "Thanks for sharing... ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 718926,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2020-01-15T01:03:22.753000",
      "content": "<p>[Update]\n<code>\nmodel: se-resnext50 with mixup + cutmix\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.988011\nLB:  0.9790\n</code>\n<code>python\nif np.random.rand()&lt;0.5: do mixup\nelse: do cutmix\n</code></p>",
      "votes": 37,
      "replies": [
        {
          "id": 719043,
          "author_name": "Ildoo Kim",
          "author_url": "",
          "post_date": "2020-01-15T04:04:42.433000",
          "content": "<p>Thanks for the update. It is really helpful.</p>\n\n<p>How do you process input image? Is gray-scaled 3x137x236 image normalized? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 719061,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-15T04:43:04.863000",
          "content": "<p>With my training loss being 0.01 and cross validation loss as 0.10, there is some overfitting. Generalisation will improve results</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 719086,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-01-15T05:49:26.897000",
          "content": "<p><a href=\"/bibek777\">@bibek777</a>  I think you have saved an average of a week or two of time for many people ;) by your wonderful posts</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 719093,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-15T05:57:11.130000",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> I get your concern. I will be careful from next time. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 719100,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-01-15T06:09:07.707000",
          "content": "<p>Sharing ideas that can reach gold zone is always controversial (but it's ok now since we have 2 month left), anyway, congratulation for you have actually reached 0.98 zone (you only need to modify image size to 224x224 for truly reaching that line)</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 719102,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-15T06:13:04.893000",
          "content": "<p>Now I know what I will do tonite🙏 🙏 🙏 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 719133,
          "author_name": "timetraveller",
          "author_url": "",
          "post_date": "2020-01-15T07:02:22.797000",
          "content": "<p><a href=\"/bibek777\">@bibek777</a> thanks for sharing. How are you able to test things out so quickly? How long is each epoch taking for you?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 719232,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-15T09:10:38.823000",
          "content": "<blockquote>\n  <p>How are you able to test things out so quickly? </p>\n</blockquote>\n\n<p>when you are in top10, you spend your time kaggling rather than timetravelling😜 😜 😜 </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 719467,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-15T14:25:54.913000",
          "content": "<p><a href=\"/timetraveller98\">@timetraveller98</a>  , I am in his team , and its hard for me to keep up ..lol :P </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 719662,
          "author_name": "Alimbekov Renat [dsmlkz]",
          "author_url": "",
          "post_date": "2020-01-15T17:54:08.770000",
          "content": "<p>Thx for sharing. Did yоu use Sampler?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 732794,
          "author_name": "Shayekh Islam",
          "author_url": "",
          "post_date": "2020-01-30T09:21:58.763000",
          "content": "<p>Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 717760,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-01-13T15:34:28.717000",
      "content": "<p>The submission of\n<code>LB 0.9875</code>\nwas a single fold model which scored on my local experiment by\n<code>CV 0.996792</code></p>\n\n<p>Hope this will be motivation for you guys to improve the performance of single model, good luck!</p>",
      "votes": 36,
      "replies": [
        {
          "id": 717781,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-13T16:02:43.820000",
          "content": "<p>Great. I'm guessing your current LB score is by ensembling! What was your baseline setup, like model architecture, optimizer or validation strategies? What about the data augmentation approach, any advantages with that? A single model is scoring really well, and I doubt whether we need to the ensemble! 🙄 </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 717786,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-01-13T16:09:21.653000",
          "content": "<p>Why not read some posts from previous image competitions like these: \n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065</a>\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107926\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107926</a>\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117210\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117210</a>\n<a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/118080\">https://www.kaggle.com/c/understanding_cloud_organization/discussion/118080</a>\n<a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/110543</a></p>\n\n<p>I'm not saying the ideas in those posts are useful for this competition too (you can check by yourself), but saying that you can learn how they were thinking during competition.</p>",
          "votes": 21,
          "replies": []
        },
        {
          "id": 717788,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-13T16:13:12.303000",
          "content": "<p>I went through this one for sure 🖤 😃 \n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107987\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107987</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 717792,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-01-13T16:19:11.490000",
          "content": "<p>😂 😂 Thank you for bringing it up.\nIt was a hard competition to me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 717793,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-13T16:20:21.107000",
          "content": "<p>ahhh APTOS, good times.... was so stressful =) </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 717797,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-01-13T16:24:08.240000",
          "content": "<p>Yea... Exactly it was\nSo let me bring up your excellent post as well 😄 \n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108030\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108030</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 718004,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-13T21:42:15.690000",
          "content": "<p>Did you balance your training data?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 723165,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2020-01-19T15:19:55.223000",
      "content": "<p>CV: 0.9913\nLB: 0.9848</p>\n\n<p>though I still haven't reached a single-model score as high as <a href=\"/haqishen\">@haqishen</a> but I haven't used any TTA and big boys like seresnext or high efficientnet/densenet due to free time and computational difficulties so I think 0.99 from single model is very achievable. Have fun single model racing guys :D</p>",
      "votes": 28,
      "replies": [
        {
          "id": 723659,
          "author_name": "Tian",
          "author_url": "",
          "post_date": "2020-01-20T09:34:31.647000",
          "content": "<p>Oh, you used magic.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 742587,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2020-02-11T11:52:46.237000",
      "content": "<p>model: se-resnext50\nimg_size: 3x137x236\naugmentation: rotate, cutmix \nCV : 0.994\nLB:  0.985</p>\n\n<p>I can not get the CV score more than 0.997 from some top kagglers. So perhaps I can get some advice and help from your guys here😃 </p>",
      "votes": 22,
      "replies": [
        {
          "id": 742607,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-02-11T12:09:05.357000",
          "content": "<p>Great work! How many epochs did you train for and what type of lr scheduler?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742611,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2020-02-11T12:11:53.787000",
          "content": "<p>80 epochs</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 742685,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2020-02-11T13:38:38.387000",
          "content": "<p>I can't surpass CV 0.982. Really curious what the top did 😲 </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 742696,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2020-02-11T13:47:04.823000",
          "content": "<p>hah.. same here😆 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742860,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-02-11T15:48:07.503000",
          "content": "<p>I have tried 300 epochs and get better lb result, not use early stopping😄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742871,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2020-02-11T15:59:55.160000",
          "content": "<p>300 epochs is too long😨 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742935,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-11T16:44:20.500000",
          "content": "<p>How did you guys handle the data imbalance? or did you just use many augmentations and train for longer? Personally I used weights in the loss function which give me better lb score but i cant get as high of CV score as you guys!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 743019,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-02-11T17:55:33.600000",
          "content": "<p><a href=\"/garybios\">@garybios</a> May I wonder, did you alter standard se-resnext architecture in any way? What loss did you use? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 743504,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-02-12T05:42:47.713000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 745944,
          "author_name": "Ang Li",
          "author_url": "",
          "post_date": "2020-02-14T12:15:08.527000",
          "content": "<p>Thanks for your sharing! Many similar things with you! Could you share which optimizer and LR do you choose?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750495,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-19T12:49:37.197000",
          "content": "<p>May I ask what ratio do you use to balance the loss between root, vowel and consonant, if this applies to you? Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754968,
          "author_name": "Manoj",
          "author_url": "",
          "post_date": "2020-02-24T09:30:10.940000",
          "content": "<p>Hi <a href=\"/garybios\">@garybios</a> , \nCan you give us an idea  about what is minimum GPU memory and run time required to train this kind of Model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759248,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-02-28T19:30:54.057000",
          "content": "<p>Did you use a pre-trained model?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 734490,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "2020-02-01T14:56:00.117000",
      "content": "<p>I met kernel error, so I couldn't submit :(\n```\n- img_size: 137x236\n- augmentation: auto augment, augmix\n- OHEM</p>\n\n<p>cv: 0.997\nLB: 0.9846\n```</p>",
      "votes": 20,
      "replies": [
        {
          "id": 734498,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-02-01T15:19:40.937000",
          "content": "<p>for people like me who didn't know about OHEM; it is Online Hard Example Mining and <a href=\"http://www.erogol.com/online-hard-example-mining-pytorch/\">here</a> is a blog explaining more about it</p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 734535,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-01T16:15:26.313000",
          "content": "<p><a href=\"/bibek777\">@bibek777</a>, I was wondering what was OHEM; was it just an expression or some kind of approach. But now I'm seeing your reply here. So, Thanks. 😅 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 734539,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-02-01T16:23:29.690000",
          "content": "<p>thats impressive</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 734592,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-01T17:37:48.300000",
          "content": "<p><a href=\"/phalanx\">@phalanx</a></p>\n\n<p>please check my starter kit, it should help you to make a submission kernel.\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757</a></p>\n\n<p>(i think you are having memory issues)</p>\n\n<p>\"cv: 0.997\", this should give you lb score of about 0.99+</p>\n\n<p>\"OHEM\" : i think this is the contribution of your score, it also confirm my (and probably other kagglers) observation:</p>\n\n<ol>\n<li>results can varies with different split</li>\n<li>some people find mixup/cutout works, other are still struggling to make it work</li>\n<li>there are only a few confusing class , most of them actually work work very well and score almost 100% in validation. further top-2 accuracy is almost 100%</li>\n<li>this competition seems to be data augmentation/sampling/synthesis problem</li>\n</ol>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 734625,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-01T18:22:53.290000",
          "content": "<p>Im very unfamiliar with OHEM, so my question is, does it kind of solve the imbalance dataset problem? Since it gives more attention to harder examples/bad performing examples. This might be a dumb question, i'm still a noob haha</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 734708,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2020-02-01T21:26:03.740000",
          "content": "<p>Thanks. Very interesting. <br>\nI suppose if OHEM is effective, Focal-Loss should be also (I do not have any idea which is better).  IMO, OHEM will be a kind of hard-encoded Focal-Loss (top-k selected).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 734774,
          "author_name": "MachineLP",
          "author_url": "",
          "post_date": "2020-02-02T01:43:59.233000",
          "content": "<p>mixup/cutmix with ohem loss : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 734804,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-02-02T02:52:16.967000",
          "content": "<p>I believe this is the original paper <a href=\"https://arxiv.org/pdf/1910.00762.pdf\">https://arxiv.org/pdf/1910.00762.pdf</a>\none advantage on focusing on top % of  biggest <code>loosers</code> is you can accelerate training =) </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 734829,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-02T04:05:49.047000",
          "content": "<p>is there a pytorch version of auto augment with the RL policy selection? I think it would take a long time to train</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 734852,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2020-02-02T05:07:45.907000",
          "content": "<p>Faster AutoAugment\n<a href=\"https://arxiv.org/abs/1911.06987\">https://arxiv.org/abs/1911.06987</a>\nExisting method use black box search algorithm, while they introduce approximate gradients to update policy and it achieve faster policy search.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1620223%2F8f918c82779294a92858014a2204d43d%2F2020-02-02%2014.06.38.png?generation=1580620035468950&amp;alt=media\" alt=\"\"></p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 734856,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2020-02-02T05:16:41.323000",
          "content": "<p>About OHEM, there are better methods than it.\nIn my experience, below 2 methods are effective for detection and classification task.\nPlease check it.</p>\n\n<p>REDUCED FOCAL LOSS\n<a href=\"https://arxiv.org/abs/1903.01347\">https://arxiv.org/abs/1903.01347</a></p>\n\n<p>Class-Balanced Loss\n<a href=\"https://arxiv.org/abs/1901.05555\">https://arxiv.org/abs/1901.05555</a></p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 734862,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-02-02T05:28:24.197000",
          "content": "<p>0.997 is a high CV for a single model.\nI'll give a try to your ideas ;)\nthanks for sharing!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 734938,
          "author_name": "MachineLP",
          "author_url": "",
          "post_date": "2020-02-02T09:01:00.883000",
          "content": "<p>thanks for sharing.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 735504,
          "author_name": "DatNT",
          "author_url": "",
          "post_date": "2020-02-03T05:58:19.860000",
          "content": "<p>from my experience, 997 local can give you ~988 - 989 lb</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 735528,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2020-02-03T06:46:36.263000",
          "content": "<p>In my experiment, intensive focal loss did not work. But Reduced Focal Loss, less intensive, looks good. We can adjust a loss curve with a threshold ( cut-off) and gamma.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F2a68cc68848650a74f541083fc1f2d5d%2Ffocal-loss.png?generation=1580711991399029&amp;alt=media\" alt=\"\"></p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 736828,
          "author_name": "DatNT",
          "author_url": "",
          "post_date": "2020-02-04T15:56:06.013000",
          "content": "<p>did not try OHEM yet but i dont have much luck with focal loss / reduced focal loss =(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 736866,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-04T16:41:50.153000",
          "content": "<p><a href=\"/moewie94\">@moewie94</a> Same thing for me, I've tried different gamma and wasn't better than my original set up</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 737066,
          "author_name": "Rohit Agarwal",
          "author_url": "",
          "post_date": "2020-02-04T21:43:52.227000",
          "content": "<p>Thanks a lot <a href=\"/bibek777\">@bibek777</a> for such a simple explanation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 756554,
      "author_name": "Uday Kamal",
      "author_url": "",
      "post_date": "2020-02-25T20:42:40.280000",
      "content": "<p>Model: Densenet121\nimg_size: 3x224x224 (simple resize)\nAugmentation: NOT cutmix or mixup\nEpoch: 40 (still running) \nCV: 0.9938\nLB: 0.9825</p>\n\n<p>To my surprise, not all augmentation methods suit all architectures. You might have to find out which augmentation will work best for your model.</p>",
      "votes": 18,
      "replies": [
        {
          "id": 756572,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-25T21:09:52.530000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756709,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-02-26T01:38:48.360000",
          "content": "<p><a href=\"/udaykamal\">@udaykamal</a> My guess is augmix? I recall that phalanx had a similar CV/LB. Shanks for sharing :P Really cool results</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 757008,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-26T10:35:21.800000",
          "content": "<p>\"To my surprise, not all augmentation methods suit all architectures. You might have to find out which augmentation will work best for your model.\"</p>\n\n<p>my experiment shows that even even for the same model, different augmentation may be required for different input image size or different training epoch. (I hope i am not over fitting my validation set)</p>\n\n<p>this prompts me to look for automatic automatic augmentation hyperparameters tuning methods</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 757361,
          "author_name": "Youhan Lee",
          "author_url": "",
          "post_date": "2020-02-26T17:16:38.637000",
          "content": "<p>Thanks for sharing. I think it's time to get a larger model with some insights from all experiments.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763293,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T10:28:02.610000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 733514,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2020-01-31T08:43:57.423000",
      "content": "<p>cv: 0.9890\nlb: 0.9794</p>\n\n<p><code>\nmodel: se_resnext50_32x4d\nimgsize: 128x128\nsplit: 5/6 train, 1/6 valid\ninference: 15 minutes (kaggle kernels)\nno tta, no ensemble\n</code></p>",
      "votes": 18,
      "replies": [
        {
          "id": 733538,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-31T09:21:47.050000",
          "content": "",
          "votes": -5,
          "replies": []
        },
        {
          "id": 733962,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-31T18:37:47.127000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 733967,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-31T18:52:36.683000",
          "content": "<p><code>tta</code> means test-time-augmentation. <code>tta</code> is a popular technique used in Kaggle to increase model performance. An example of tta would be to apply horizontal flips to test images and make prediction on flipped image, and then average with the usual image. Below is a pseudo-code for <code>tta</code>. Hope this helps you to understand the concept of <code>tta</code></p>\n\n<p>```python\nimage = get_image(img_name)\npredict_1 = model(image)</p>\n\n<p>flip_image = apply_hflip(image)\npredict_2 = model(flip_image)</p>\n\n<p>final_pred = take_mean(predict_1, predict_2)\n```</p>",
          "votes": 16,
          "replies": []
        },
        {
          "id": 734003,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-01-31T19:59:32.023000",
          "content": "<p>Nice work! How many epochs did you train for?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 734194,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2020-02-01T04:22:46.143000",
          "content": "<p>40 epochs.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 734263,
          "author_name": "kaerururu",
          "author_url": "",
          "post_date": "2020-02-01T07:17:52.067000",
          "content": "<p>Thank you for your information!\nDid you use imagenet weight ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 734270,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2020-02-01T07:29:13.013000",
          "content": "<p>Hi kaeru-san, \nYes, I use imagenet weights from pytorch <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">pretrainedmodels</a>. I haven't tried it without imagenet weights because I believe it should work better than random weights at least.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 734413,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-02-01T12:47:33.480000",
          "content": "<p>Hi <a href=\"/appian\">@appian</a> wonderful results! I am nowhere near that with single model. If you don't mind disclosing, what is the batch size you used for training? Number of epochs is not an effective parameter without the batch size. Thx in advance</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 734488,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-01T14:48:36.790000",
          "content": "<p>I'm using the same model and found that without pretrained weights it performs way worse</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 735966,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2020-02-03T17:11:23.760000",
          "content": "<p><a href=\"/appian\">@appian</a> do you use any particular trick for convergence? It's been very hard for me to reach 0.97 on validation set even after ~50 epochs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 737097,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2020-02-04T23:06:54.993000",
          "content": "<p><a href=\"/tahsin\">@tahsin</a> \nNot really. I just use Adam with reducelronplateau. The score is achievable with ideas discussed in this competition so far. Qishen has summarized these ideas and I think it's helpful.\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/127976\">https://www.kaggle.com/c/bengaliai-cv19/discussion/127976</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 737825,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-05T20:19:37.780000",
          "content": "<p><a href=\"/appian\">@appian</a> For reducelronplateau do you monitor val loss or CV?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 741179,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2020-02-10T09:49:47.130000",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>\nI monitored val loss with patience of 5. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 741205,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-10T10:15:15.917000",
          "content": "<p><a href=\"/appian\">@appian</a> \nwould please inform what was the lowest validation loss you have got? I also monitor validation loss, my lowest validation loss was around 0.21 something. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 742328,
          "author_name": "kaerururu",
          "author_url": "",
          "post_date": "2020-02-11T08:36:08.560000",
          "content": "<p><a href=\"/appian\">@appian</a> \nThank you for your replying !!\n(I forgot my asking you ... very sorry m(_ _)m)</p>\n\n<p>Next experiments I'll use imagenet weight !!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 743912,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2020-02-12T11:47:56.603000",
          "content": "<p><a href=\"/kaerunantoka\">@kaerunantoka</a>\nNo worries. Good luck with experiments!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 746084,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2020-02-14T15:49:38.210000",
          "content": "<p>Hi <a href=\"/appian\">@appian</a>, what augmentation techniques are you using? Mixup, Cutmix, Gridmask, augmix? or combination of them ?  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 751641,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2020-02-20T11:10:43.277000",
          "content": "<p>UPDATE</p>\n\n<p>CV: 0.9982\nLB: 0.9885\nepoch: 70</p>\n\n<p>single model, no tta, no ensemble. </p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 751762,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-02-20T13:47:36.507000",
          "content": "<p>Dude, you're killing it. Congrats</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 751781,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-20T14:05:13.420000",
          "content": "<p>May I ask if you use mixup/cutmix?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 762424,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2020-03-03T13:54:58.717000",
      "content": "<p><code>update</code>\ncv: 0.9981\nlb: 0.9900\nsingle fold, no tta\nI still haven't found <a href=\"https://www.kaggle.com/haqishen\">haqishen‘s</a> magic😑 😑 </p>",
      "votes": 15,
      "replies": [
        {
          "id": 762431,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-03-03T14:01:51.453000",
          "content": "<p>single fold! awesome.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 762446,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-03-03T14:15:16.287000",
          "content": "<p>Yes! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762448,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-03-03T14:17:44.557000",
          "content": "<p>wow</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762471,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-03-03T14:36:24.767000",
          "content": "<blockquote>\n  <p>I still haven't found haqishen‘s magic😑 😑</p>\n</blockquote>\n\n<p>so you created your own 👍 👍 </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 762693,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-03-03T18:14:15.877000",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4059244%2Faca408f0846c7509bb31f552dac21b53%2FCapture.PNG?generation=1583259244730559&amp;alt=media\" alt=\"\"></p>",
          "votes": 15,
          "replies": []
        },
        {
          "id": 762745,
          "author_name": "Morphy",
          "author_url": "",
          "post_date": "2020-03-03T19:22:29.333000",
          "content": "<p>Gary NB!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 769780,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-03-12T09:05:20.870000",
          "content": "<p>666</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 735506,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2020-02-03T06:01:24.957000",
      "content": "<p>local: <code>0.9970</code>\nlb: <code>0.9884</code></p>\n\n<p>single model, single fold, no tta</p>\n\n<p>i still think <code>0.99</code> is really doable but the you have to use big models/ tta and maybe some lucky with random seed because improvement seems quite random when local error is small</p>",
      "votes": 15,
      "replies": [
        {
          "id": 735521,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-03T06:33:58.553000",
          "content": "<p>wow, that' huge for single model. However I don't understand how it's so doable! I can't score minimum .97. Its seem some people are easily getting high score :( </p>\n\n<p>My basic setup: \n(keras)</p>\n\n<p><code>\nmodel: efficient \nimg size: 128\nsplit: 80/20\nopt: adam\naugmentation: augmix \nepoch: 20\n</code>\nImplementing almost same result, some people are getting really promising score.  :(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 736840,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-04T16:12:16.893000",
          "content": "<p>Personally i haven't been having good results with efficient net maybe you should experiment with other architectures! :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 737672,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-05T16:28:26.187000",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> are you using keras? If you don't mind, if you are, would you please inform the best outcome of your findings with efficientnet? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737677,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-02-05T16:32:54.900000",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> If you are using augmentation like mixup/cutmix/cutout/augmix/etc. you have to train your model much longer. try it for 100-150 (or more) epochs. After 20 epochs one of my model's (cutmix) CV was ~0.968; after 120 epochs it was 0.9815</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 737687,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-05T16:46:40.953000",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> thank you for your tips. Actually I went for 200 epoch but gave an early stop for validation loss at epoch 30. I mean, if the validation loss didn't improve within 30 epoch, the training process would stop. And my model stopped at around epoch 31~33. I am not sure, should I increase the patience epoch size which is currently 30 or just pick the last optimized weights?</p>\n\n<p>Again, in this competition, an open secret is to use <code>cutmix-mixup</code> to get high score but It's comparatively hard for me to implement it right now, I am using Keras. I am currently using augmix-cutout-gridmask. Would you please inform, what is your individual validation score of the three target: <code>grapheme_root</code> , <code>vowel</code> and <code>consonant</code> ? I found that score almost  .98/.99 in both vowel and consonant is doable but score above .96/.97 for grapheme_root is pretty tough. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 737694,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-02-05T16:53:39.300000",
          "content": "<p>you should implement official metric for this competition and use this as as early stop. I am very confident this will improve your results. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 737704,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-02-05T17:05:51.913000",
          "content": "<p>Grapheme root: 0.9752, vowel: 0.9855, consonant: 0.9924\nI suggest (at least for one trial) remove early stopping. </p>\n\n<p>Here is a screenshot from my training, it might help. (open it in a new tab for higher resolution.)</p>\n\n<p><img src=\"https://albumizr.com/ia/8059658e005e0c58ed268812c1b1f7d1.jpg\" alt=\"\"></p>\n\n<p>Edit:\nI use iterations instead of epochs (every datapoint/validation step is after 1000 iterations/forward passes; batch: 64; valid size: 8%; ~2887 iterations = one epoch)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 737770,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-05T18:53:35.607000",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> Well im using a xresnet34 which i know will not give me the best results but i have computational limitations.. i added 3 tails just like <a href=\"/drhabib\">@drhabib</a> experimented in his post: \n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123432\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123432</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 739024,
          "author_name": "Tushar",
          "author_url": "",
          "post_date": "2020-02-07T09:51:07.260000",
          "content": "<p>Which tool are you using <a href=\"/pestipeti\">@pestipeti</a> for these visualizations? It looks amazing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 739029,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-02-07T10:02:35.383000",
          "content": "<p><a href=\"/thanatoz\">@thanatoz</a> I use <a href=\"https://neptune.ai\">neptune.ai</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 739132,
          "author_name": "Eek The Cat",
          "author_url": "",
          "post_date": "2020-02-07T12:47:38.493000",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> may I ask, when you trained a model like this \" try it for 100-150 (or more) epochs. After 20 epochs one of my model's (cutmix) CV was ~0.968; after 120 epochs it was 0.9815\", did you use any LR scheduling? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 739150,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-02-07T13:22:53.533000",
          "content": "<p><a href=\"/cateek\">@cateek</a> Yes, I used CosineAnnealingLR</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 741323,
          "author_name": "Mayank Pathak",
          "author_url": "",
          "post_date": "2020-02-10T13:35:45.003000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  whats the good learning rate and optimizer for this problem ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 741526,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-02-10T18:44:11.137000",
          "content": "<p><a href=\"/mayank17\">@mayank17</a> I haven't done so much systematic studies with optimizers. So far <code>Adam</code> with LR of <code>0.003</code> and <code>ReduceOnPlatau</code> or <code>OneCycleLearning</code> policies works fine. Generally its best to stick with one optimizer and once you are satisfied with your model performance you can play around =) I will update once I have results.</p>\n\n<p>Good luck </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 747639,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-02-16T17:22:57.523000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753198,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-02-21T20:47:14.013000",
          "content": "<p>mäh</p>",
          "votes": -3,
          "replies": []
        },
        {
          "id": 757578,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-26T23:33:15.720000",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Hi Peter May I ask if you pre-process the image, i.e. centered and cropped or just resize the raw image.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 757593,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-02-27T00:00:43.193000",
          "content": "<p><a href=\"/yl1202\">@yl1202</a>, for my experiment above I preprocessed the image (I used <a href=\"/iafoss\">@iafoss</a> kernel). For my current best (0.9799) I only resized the images to 128x128px. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 757605,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-27T00:14:32.253000",
          "content": "<p>Thanks Peter. This is really interesting...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 761883,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2020-03-03T02:28:32.213000",
      "content": "<p>You guys may believe or not. \nYesterday, I dreamed about a solution of the top team. Even I could not remember the details, however, when I waked up, I had an idea to follow up. It gave me a major improvement 😂  </p>\n\n<p>img_size: 137x236\nCV: 0.993 \nLB: 0.9868 \nSingle fold, no tta. </p>",
      "votes": 15,
      "replies": [
        {
          "id": 761884,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2020-03-03T02:29:27.670000",
          "content": "<p>I am going to sleep now. I hope to see other top's solutions 😆 </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 761890,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-03-03T02:36:20.387000",
          "content": "<p>we probably saw the same dream ... but i remembered a little bit more...</p>\n\n<p>single model\n<code>\ncv: 0.9975\nlb: 0.9893\n</code></p>\n\n<p>single fold .. no TTA =)</p>\n\n<p>Also credit goes to my team =) </p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 761895,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-03-03T02:45:47.340000",
          "content": "<p>I need that dream right about now... :P</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 761896,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2020-03-03T02:45:48.017000",
          "content": "<p>I guess the solution I dreamed about is yours. :))</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 761979,
          "author_name": "Helen",
          "author_url": "",
          "post_date": "2020-03-03T05:13:33.007000",
          "content": "<p>Couldn't wait to dream your dream😄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762006,
          "author_name": "Youhan Lee",
          "author_url": "",
          "post_date": "2020-03-03T05:35:52.713000",
          "content": "<p>Also credit goes to my team =)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 762461,
          "author_name": "Mohammad Azam Khan",
          "author_url": "",
          "post_date": "2020-03-03T14:31:35.267000",
          "content": "<p>What a sweet dream <a href=\"/backaggle\">@backaggle</a> and <a href=\"/drhabib\">@drhabib</a>!\nBTW, did you guys dream of any augmentation at that time 👀 ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 762470,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-03-03T14:36:21.050000",
          "content": "<p>yes I dont remember everything but it was something cutting and mixing .... perhaps cutmix ... ? and some images were rotated.... not sure.... everything is so cloudy... </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 763268,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-03-04T10:06:04.887000",
          "content": "<p><a href=\"/backaggle\">@backaggle</a> few days ago i saw a dream which  was \"we trained 2 models and 1 of them was trained for grapheme roots and by combining those 2 models (1 for full data and 1 for grapheme root) gave us a huge boost,i remember full dream \"i saw we are very close to gold zone after executing this plan,then didn't give it a try ha ha ha\nyou had a sweet dream :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 763687,
          "author_name": "Gold Retriever",
          "author_url": "",
          "post_date": "2020-03-04T18:37:17.063000",
          "content": "<p>I think it's time for some Inception... @christophernolan</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 738975,
      "author_name": "ccchang",
      "author_url": "",
      "post_date": "2020-02-07T08:35:04.517000",
      "content": "<p>Model: se-resnext50-32x4d\nImage size: 128x128x1\nCV: 0.9937\nLB: 0.9838</p>\n\n<p>Interesting that no matter how I changed model structure, augmentation or image size,\nLB scores are always equal to my CV scores minus about 1~1.3%, \nguess I need totally different way to break through 99%</p>",
      "votes": 13,
      "replies": [
        {
          "id": 738980,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2020-02-07T08:40:54.857000",
          "content": "<p><a href=\"/ccchang801023\">@ccchang801023</a> , how many epochs are you training for?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738981,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-07T08:40:59.203000",
          "content": "<p>Would you please inform what is your individual score of the three target output? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 739020,
          "author_name": "Tushar",
          "author_url": "",
          "post_date": "2020-02-07T09:42:28.423000",
          "content": "<p>And you are achieving this without Mixup/Cutmix?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 739233,
          "author_name": "ccchang",
          "author_url": "",
          "post_date": "2020-02-07T15:28:36.203000",
          "content": "<p><a href=\"/pheadrus\">@pheadrus</a> 150 epochs\n<a href=\"/ipythonx\">@ipythonx</a>  My CV :  0.5 * 0.99107(root) + 0.25 * 0.99648(vowel) + 0.25 * 0.99641(consonant)  = 0.9937\n<a href=\"/thanatoz\">@thanatoz</a>  Yes I combined augmentation methods and Cutmix is one of them</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 739272,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-07T16:18:25.830000",
          "content": "<p><a href=\"/ccchang801023\">@ccchang801023</a> are those co-efficient (.5, .25, .25) loss weights? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 739283,
          "author_name": "ccchang",
          "author_url": "",
          "post_date": "2020-02-07T16:26:32.023000",
          "content": "<p>No, just the weights described in evaluation metric : final_score = np.average(scores, weights=[2,1,1])</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 739285,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-07T16:36:54.043000",
          "content": "<p>I must say score of root is pretty impressive of yours. Have you addressed any data imbalance for it particularly or just lots of augmentation process and hard training with care? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 745943,
          "author_name": "Ang Li",
          "author_url": "",
          "post_date": "2020-02-14T12:13:18.277000",
          "content": "<p>Which optimizer and LR do you choose?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 737631,
      "author_name": "Inoichan",
      "author_url": "",
      "post_date": "2020-02-05T15:50:09.627000",
      "content": "<p><strong>CV : 0.979</strong>\n*<em>LB : 0.971</em>*\n<code>\nModel: SEResNeXT50\nAugmentations: 40% Mixup + 40% Cutmix + 10% Cutout + 10% GridMask\nSplit: Character stratified split 80/20\nImage size: 128x128x3 (just resize)\nEpoch: 100\nSingle fold, No TTA\n</code></p>\n\n<p>[Next]\n- AugMix\n- Reduced Focal Loss\n- Bigger image size\n...</p>\n\n<p>There are lot of things to be done before getting to 0.99...</p>\n\n<p>I would appreciate if you could give me any advice!!</p>",
      "votes": 14,
      "replies": [
        {
          "id": 737636,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2020-02-05T15:56:12.580000",
          "content": "<p>Appreciate you posting details. Quick question Character stratified split is equivalent to iterative stratified splits? (<a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a>)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737670,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-05T16:25:15.550000",
          "content": "<p>Impressive. Few query:\n- have you used <a href=\"/iafoss\">@iafoss</a> 's processing method?\n- epoch: 100; have you used early stop or just full 100 epoch and pick up the last updated weights?\n- about the augmentation split, would you please inform how to achieve this splitting method, I mean the percentage?\n- choosing stratified split, (its a general question) why you went only single fold? In cross validation, if you made 5 fold: 1 for test, 4 for train and you train the model. 5 fold will give you 5 times different trained model, won't it? Would you please clarify? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738721,
          "author_name": "Gold Retriever",
          "author_url": "",
          "post_date": "2020-02-06T22:54:53.467000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738989,
          "author_name": "Inoichan",
          "author_url": "",
          "post_date": "2020-02-07T08:57:52.757000",
          "content": "<p>Sorry for the late reply..</p>\n\n<p><a href=\"/pheadrus\">@pheadrus</a> \nCharacter stratified split may be different from  iterative stratified splits. I'm using stratified split by 1295 graphems (Bengali characters). This is because at first I have tried 1295 class classification model,  or custom models like a following figure. However, these 1295 class models were tend to overfit than 3 outputs models... Further experimentation is needed. My LB / PB score above is just a 3 outputs model.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1057275%2Fd267cf55d28be8a24e39f3ceb2534def%2F_cut.png?generation=1581063579323056&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"/ipythonx\">@ipythonx</a> </p>\n\n<blockquote>\n  <p>have you used <a href=\"/iafoss\">@iafoss</a> 's processing method?</p>\n</blockquote>\n\n<p>I haven't used it yet. I'll try it!!</p>\n\n<blockquote>\n  <p>epoch: 100; have you used early stop or just full 100 epoch and pick up the last updated weights?</p>\n</blockquote>\n\n<p>I didn't use early stopping. Further speaking, I think, 100 epochs are still not enough.</p>\n\n<blockquote>\n  <p>about the augmentation split, would you please inform how to achieve this splitting method, I mean the percentage?</p>\n</blockquote>\n\n<p>I experimented mixup only training, cutmix only training, and so on. I decided the ratio by the result, that is, mixup and cutmix were equally good, and gridmask and cutout were a little worse than mixup and cutmix. About the ratio, I need to experiment more.</p>\n\n<blockquote>\n  <p>choosing stratified split, (its a general question) why you went only single fold? In cross validation, if you made 5 fold: 1 for test, 4 for train and you train the model. 5 fold will give you 5 times different trained model, won't it? Would you please clarify?</p>\n</blockquote>\n\n<p>You are right. But, even training the model once takes a very long time. So, after I find the better method, I will finally train the model 5 times as you said. </p>\n\n<p>(Let me apologize for my poor English. )</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 739256,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2020-02-07T16:02:32.180000",
          "content": "<p>How are you weighing the 4 losses? I think its interesting to also consider 1295 class classification. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 740277,
          "author_name": "Inoichan",
          "author_url": "",
          "post_date": "2020-02-09T06:53:46.673000",
          "content": "<p><a href=\"/pheadrus\">@pheadrus</a> \nAs shown below. (with results)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1057275%2F30898817da60d9f7500f6869bc759211%2F.png?generation=1581231174993166&amp;alt=media\" alt=\"\"></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 744362,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-12T19:10:35.607000",
          "content": "<p>What program did you use to create those figures (if you used one)? <a href=\"/inoueu1\">@inoueu1</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 745006,
          "author_name": "Inoichan",
          "author_url": "",
          "post_date": "2020-02-13T11:40:39.730000",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> \nI'm using google drawings👍 </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 766744,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-03-08T16:21:20.507000",
      "content": "<p><code>\nSingle Fold:\nCV: 0.9967\nLB: 0.9884 \n</code></p>\n\n<p>EfficientNetB4, 130x224, cutmix, OHEM </p>",
      "votes": 11,
      "replies": [
        {
          "id": 766764,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-08T16:51:06.747000",
          "content": "<p>any comparison results of OHEM with focal loss?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 766808,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-08T19:06:29.470000",
          "content": "<p>what is OHEM？</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 766848,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-03-08T20:48:04.280000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> I didn't try focal loss so can't comment on that. OHEM was working well for me, so I just stuck with it. </p>\n\n<p><a href=\"/xiaolonglee\">@xiaolonglee</a> OHEM is online hard example mining. Essentially, it only backpropagates the largest losses (tunable) in each minibatch when training, the idea being that easy examples with low loss don't really contribute to learning. See here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 766851,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-08T20:58:27.780000",
          "content": "<p>Thank you for your guidence !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 766937,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-09T00:36:05.297000",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> \nthank you for the reply.</p>\n\n<p>i am thinking that focal loss should work, but i cannot get good results with it. this puzzles me. do you have an estimate of the contributions of each step?</p>\n\n<p>in my experiments, i am having:</p>\n\n<p>EfficientNetB4, 137x236 (no augmentation) : 0.974 at local cv\nadd basic augmentation (scale, rotate, shift) : 0.978\nadd cutmix : 0.992\nadd focal loss : 0.992</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 766961,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-03-09T01:30:24.447000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> </p>\n\n<p>Baseline EfficientNet-B4:  0.978\nAdd cutmix: 0.986\nAdd OHEM: 0.997</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 767019,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-03-09T03:28:49.453000",
          "content": "<p>Wow, Ohem loss gave you a big boost. We found it made our se-resnext worse. Will definitely have to try on effnet. Thanks Ian</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767048,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-03-09T04:58:16.263000",
          "content": "<p>For how many epochs did you train? Is training for 100+ epochs a big factor for your result?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767056,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2020-03-09T05:15:04.973000",
          "content": "<p>Hi <a href=\"/vaillant\">@vaillant</a> , can u pls elaborate on when did u apply OHEM during training? Did u apply OHEM from the beginning of training itself or u switched to it in the middle of training like after 50 epochs or so? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767277,
          "author_name": "Shayekh Islam",
          "author_url": "",
          "post_date": "2020-03-09T12:17:37.573000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> How to select EfficientNet backbone, i mean b4 or b5?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767304,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-03-09T13:12:21.543000",
          "content": "<p><a href=\"/kurianbenoy\">@kurianbenoy</a> 60 epochs. I tried 120 and didn't see any difference.\n<a href=\"/virajbagal\">@virajbagal</a> I don't use it at the beginning. \n<a href=\"/shayekh\">@shayekh</a> I experiment and see what works best. I favor smaller models if the performance is essentially the same. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 767671,
          "author_name": "ratan rohith",
          "author_url": "",
          "post_date": "2020-03-10T01:28:30.690000",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> Ohem loss gave you a good boost! Thanks for sharing.One question from my side other than cutmix are you using any augmentations? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 767687,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-03-10T02:10:34.990000",
          "content": "<p><a href=\"/ratan123\">@ratan123</a> Nope, cutmix only.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 767776,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-10T04:45:39.053000",
          "content": "<p>\"How to select EfficientNet backbone, i mean b4 or b5?\" </p>\n\n<p>see how much computation power you have.  b3,4,5 have similar performance</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 767939,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-10T09:24:34.420000",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> \n\"Baseline EfficientNet-B4: 0.978\"</p>\n\n<p>i am now trying to repeat your experiment. This is my experiment results that you may be interested in:</p>\n\n<p>Baseline EfficientNet-B3 (no augmentation #1 ) + ohem: cv 0.988 ~ 0.990\n  - size = 128x224\n  - change first convolution stride=2 to stride=1  </p>\n\n<p>[# 1]  i ran many experiments and had used previous trained models to for initialization for new training. my learning rate started at 0.05 with batch_size=256. it is possible that the final results is due to previously models (i.e. effects of initialization). But nevertheless, the results reported for  cv 0.988 ~ 0.990 is for training no augmentation for ~20 epochs after initialization</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 769120,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-03-11T14:47:34.460000",
          "content": "<p>OHEM Loss made my CV worse. I start without using OHEM, then slowly start increasing OHEM's effects as the epochs rise. Sadly did not help me in seresnext or efficientnet =)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 770584,
          "author_name": "Mohammad Azam Khan",
          "author_url": "",
          "post_date": "2020-03-13T05:52:47.540000",
          "content": "<p><a href=\"/vaillant\">@vaillant</a> While OHEM was not performed well for most of the participants, you got a significant performance boost. If you don't mind, by what ratio did you use for OHEM wrt batch size?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 770752,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-03-13T11:04:49.133000",
          "content": "<p><a href=\"/muhammedazamkhan\">@muhammedazamkhan</a> I gradually decrease the ratio from 1 to 1/8 throughout training. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 771034,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2020-03-13T17:23:55.497000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 771037,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-03-13T17:29:22.940000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 771433,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-03-14T07:01:49.803000",
          "content": "<p>CV 0.994  with ohem</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 761824,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "2020-03-03T00:54:56.403000",
      "content": "<p>img_size: 137x236\ncv: 0.996\nlb: 0.9894\nsingle fold, no tta</p>",
      "votes": 12,
      "replies": [
        {
          "id": 761829,
          "author_name": "Youhan Lee",
          "author_url": "",
          "post_date": "2020-03-03T01:07:23.337000",
          "content": "<p>You made quantum jump :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 761831,
          "author_name": "Karol Zak",
          "author_url": "",
          "post_date": "2020-03-03T01:11:59.907000",
          "content": "<p>Nice! Congrats! Care to share which model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762021,
          "author_name": "Bet4Honor",
          "author_url": "",
          "post_date": "2020-03-03T05:55:57.077000",
          "content": "<p>good job. My cv: 0.996, but lb only get 0.982</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 772733,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2020-03-15T20:40:11.863000",
      "content": "<p>Finally reached cv0.9975, but it's too late...\nmodel: efficientnet-b3\nsplit: 5fold with stratified k-fold\nimage size: 224x224\naugment: gridmask, cutout</p>",
      "votes": 9,
      "replies": [
        {
          "id": 773819,
          "author_name": "Ryunosuke Ishizaki",
          "author_url": "",
          "post_date": "2020-03-16T00:07:45.573000",
          "content": "<p>Nothing is too late man ;)\nLet's enjoy our last journey Lol</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 775269,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-03-16T13:19:28.393000",
          "content": "<p>The best CV I can get is 0.991...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 775281,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-03-16T13:46:27.423000",
          "content": "<p><a href=\"/ryunosukeishizaki\">@ryunosukeishizaki</a> Thanks, I'm still training models to ensemble!!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 775282,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-03-16T13:49:16.190000",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> Higher cv doesn't mean higher private lb:) By the way my score is on TPU and using your gridmask implementation! Thanks!!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 775287,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-03-16T13:59:29.523000",
          "content": "<p><a href=\"/bamps53\">@bamps53</a> No problem! Good luck in private score!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 775331,
          "author_name": "Vishal Sharma",
          "author_url": "",
          "post_date": "2020-03-16T15:20:57.437000",
          "content": "<p>how many epochs?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 775459,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-03-16T17:44:31.580000",
          "content": "<p><a href=\"/vishal1310\">@vishal1310</a> \nAbout 150 epochs with reduce plateau with patience=10.\nGood luck, we only have 6 hours!!</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 775686,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-03-17T00:02:37.533000",
          "content": "<p>I can see why our private scores are bad, you guys must be using one head model like see-- shared, right?\nI did notice that using one head might narrow down the prediction categorical range of data, which means models were overfitting to training and public testing data. But I did't listen to my brain, since one head model just easier to train, haha. This was a great experience!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 775693,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-03-17T00:08:21.313000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2Fbb5b7423173c34fb687a1b927bc5059d%2F2020-03-17%208.05.48.png?generation=1584403618844025&amp;alt=media\" alt=\"My scores with 3 heads output\"></p>\n\n<p>My score with 3 heads output</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 775714,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-03-17T00:21:47.547000",
          "content": "<p>You're absolutely right!!!\nAt first I tried to predict with 3 head, but somehow I couldn't resolve Submission Scoring Error....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 775722,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-03-17T00:27:19.263000",
          "content": "<p>3 heads model need more time to train and inference, so you might need to optimize the submission notebook. I took a lots of time to optimize the submission notebook, but I went to chose the one head model in the end . Well, maybe the luck will appear in next competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 775958,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-03-17T03:40:16.187000",
          "content": "<p>I took a test on my final submitted notebook.\nTotally same parameter/augmentation/model but change one head to three heads(single fold): </p>\n\n<p>1 head:\nprivate score : 0.9088\npublic score : 0.9853</p>\n\n<p>3 heads:\nprivate score : 0.9348\npublic score : 0.9797</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 721664,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-17T15:32:25.553000",
      "content": "<p>I completed initial analysis using cam class activation map. Indeed there is much overfitting. Training loss can be 0.01 and validation loss can be 0.20 to 0.13. With regularisation like mixup and cutout, the cam is vastly improved.  </p>\n\n<p>Hence to get good results, one can rely on such regularization or add hand more label signal ( e.g pixel or box annotations) to improve the cam response</p>",
      "votes": 10,
      "replies": [
        {
          "id": 721767,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-17T17:12:12.487000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 716280,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2020-01-11T14:03:05.757000",
      "content": "<p>[Update]\n<code>\nmodel: se-resnext50 with mixup\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.982066\nLB:  0.9739\n</code></p>",
      "votes": 10,
      "replies": [
        {
          "id": 716313,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-01-11T14:39:20.353000",
          "content": "<p>Great result <a href=\"/bibek777\">@bibek777</a> !\nFor how many epochs did you train? Your LB score is a result of a single fold?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 716315,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-11T14:41:03.483000",
          "content": "<p>I trained it for 60 epochs. Yes, it is a single model score</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 716402,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-11T16:57:07.757000",
          "content": "<p>the difference between CV and LB after ensemble could be about 0.005.\nSince your CV is 0.982066, you can get better results after ensemble.</p>\n\n<p>(This also means that for kagglers in the rage of LB  0.985, their CV could be in the range of 0.99?)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 716424,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-11T17:38:29.060000",
          "content": "<p>Did you do mix up training like you said you were trying to figure out or just normal augmentations?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716476,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-11T18:51:57.950000",
          "content": "<p>yes, surely ensemble will give better score but how far can we get with a single model? At first, I thought <code>0.970</code> but to my surprise, so far the best is <code>0.9739</code>. I wonder if we can go to <code>0.98</code> with single model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716477,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-11T18:52:45.897000",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> I added mixup(+a lot of other tricks) for my current score</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 716486,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-11T19:03:03.207000",
          "content": "<p>Wow good job my friend, youre are a beast!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716624,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-12T02:38:46.750000",
          "content": "<p>\"I wonder if we can go to 0.98 with single model?\"</p>\n\n<p>simply change 137x236 to 224x224 using F.interpolate() at the input. you may be able to see some improvement.</p>\n\n<p>control your sampling as well, e.g. mix samples of different class , mix samples of weaker class, mix samples of same class ??? mix by root, constant or vowel?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 716702,
          "author_name": "Ildoo Kim",
          "author_url": "",
          "post_date": "2020-01-12T05:38:14.257000",
          "content": "<p>Interesting result. What is your loss function?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716772,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-12T08:26:43.600000",
          "content": "<p>It is based on <a href=\"https://github.com/hysts/pytorch_image_classification/blob/master/augmentations/mixup.py\">this repo</a> </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 716782,
          "author_name": "timetraveller",
          "author_url": "",
          "post_date": "2020-01-12T08:36:03.680000",
          "content": "<p>60 epochs for models that deep?!\nFor me, even after optimizing the data loader as much as I can, it's gonna take ~30 hours. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716952,
          "author_name": "Andrey Zotov",
          "author_url": "",
          "post_date": "2020-01-12T14:02:52.077000",
          "content": "<p>Therefore, you need to unite in a team of 5 people. And train the network in parallel. (each on his own fold, for example)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 717716,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-13T14:13:37.717000",
          "content": "<p><a href=\"/bibek777\">@bibek777</a>  congratulation on becoming discussion Master =) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 720546,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-01-16T14:15:50.747000",
          "content": "<p>What GPU do you have?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 720550,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-16T14:16:42.113000",
          "content": "<p>TitanX</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 703644,
      "author_name": "Maxwell",
      "author_url": "",
      "post_date": "2019-12-26T12:23:29.420000",
      "content": "<p>model: ResNet 18 (pretrained=False)\nimg: 1 channel </p>\n\n<p>CV: 0.9654\nLB:  0.9632</p>\n\n<p>I think CV and LB are correlated well, according to some posts in this thread :-)</p>\n\n<hr>\n\n<p>updated on Jan.02 2020</p>\n\n<p>ResNet 18\nCV: 0.9729\nLB: 0.9681</p>",
      "votes": 10,
      "replies": [
        {
          "id": 703668,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2019-12-26T13:25:54.943000",
          "content": "<p>Great. Few question though;</p>\n\n<ul>\n<li>what other models you've tried except resnet? </li>\n<li>cv or cross validation, right? how did you do that? I mean, have you made n times fold (of the whole train set) and train the model in that way? I know about cross validation process, but I am not sure why or how some people are doing that? Isn't it expensive for training?</li>\n<li>like others, the output are the tree labels..and using softmax for each label independently you score the probabilities, right? Have you tried to implement sigmoid on the last layer for three target variables at a time?</li>\n</ul>\n\n<p>Thank you.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 703690,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-12-26T14:02:08.387000",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> </p>\n\n<blockquote>\n  <p>what other models you've tried except resnet?</p>\n</blockquote>\n\n<p>Not yet now. I have just made a baseline model :-)  </p>\n\n<hr>\n\n<blockquote>\n  <p>cv or cross validation, right? how did you do that? I mean, have you made n times fold (of the whole train set) and train the model in that way? I know about cross validation process, but I am not sure why or how some people are doing that? Isn't it expensive for training?</p>\n</blockquote>\n\n<p>I used the <code>CV</code> in the meaning of Cross Validation with 5-fold. As you mentioned, cross validation process is expensive. But it will make our model evaluation robust, because we use all train data to evaluate with less leakage (Unfortunately early stopping and other hyper parameter setting can not be completely free of leakage). The process is following,</p>\n\n<p>step1. imagine you split all train data into  5 subsets(We will call these as 1, 2, 3, 4, 5) <br>\nstep2. train your model (ResNet, EfficientNet, others) with 1+2+3+4, and make inference on 5\nstep3. Do the same process as step2 (train with 2+3+4+5 and infer on 5, ...) \nstep4. Then you can get OOF(Out Of Fold) prediction on all train data and the score based on it, that's what I called CV score.</p>\n\n<hr>\n\n<blockquote>\n  <p>like others, the output are the tree labels..and using softmax for each label independently you score the probabilities, right? Have you tried to implement sigmoid on the last layer for three target variables at a time?</p>\n</blockquote>\n\n<p>Yes. I used softmax. But using sigmoid and other special structure and loss function will be alternatives.  We must do some experiments :-)</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 703694,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2019-12-26T14:07:39.467000",
          "content": "<p>Awesome. Thank you. :-)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 703930,
          "author_name": "Hanjoon Choe",
          "author_url": "",
          "post_date": "2019-12-26T20:33:08.480000",
          "content": "<p>I’ve used all train set for training achieving +0.97, and LB was 0.94. Maybe overfitting.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 704090,
          "author_name": "Dhananjay Raut",
          "author_url": "",
          "post_date": "2019-12-27T03:31:53.037000",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a> How much time did your inference kernel took? \nwith 5 Fold I suspect mine will overshoot the time limit.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 704116,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-12-27T04:28:35.970000",
          "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> </p>\n\n<p>It will take less than 30 minutes.\nAs you know size of test data is only 12, so I simulated inference time with train data.</p>\n\n<p>One useful topic was posted, <br>\n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/122993\">https://www.kaggle.com/c/bengaliai-cv19/discussion/122993</a>\nand I think others will be.</p>\n\n<p>When we ensemble models, short inference time is desirable to blend more models. Let's take care of inference time ;-)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 704190,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2019-12-27T06:32:57.340000",
          "content": "<p><a href=\"/hanjoonchoe\">@hanjoonchoe</a>  maybe too many epochs causing overfit?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 704263,
          "author_name": "Hanjoon Choe",
          "author_url": "",
          "post_date": "2019-12-27T08:27:14.257000",
          "content": "<p><a href=\"/p4rallax\">@p4rallax</a> I don't know yet. I separated train set 8:2 now, and checking two metric score. There are a bit of gap between train and valid set as opposed to the other kagglers reporting here. Oh… Maybe They only check metrics for valid only. My bad XD. I just train my model with down sampled images so far, but using whole data set at once is more effective when I see. I do not have gpu quota to inference my model to check the score now though.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 704343,
          "author_name": "Hanjoon Choe",
          "author_url": "",
          "post_date": "2019-12-27T11:00:22.927000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2791153%2Fe148f19e546d3c50c540ea24e079db8e%2FScreen%20Shot%202019-12-27%20at%206.00.02%20AM.png?generation=1577444419066240&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 707343,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2019-12-31T17:10:09.910000",
          "content": "<p><a href=\"/maxwell110\">@maxwell110</a>   Can u please describe ur training hyperparameters . I mean what was ur lr, decay and scheduler  ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715487,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-10T15:28:09.890000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 766728,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-03-08T16:05:32.770000",
      "content": "<p>My score on <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">new validation setting</a> is like this:</p>\n\n<p><code>\nLocal Score: 0.9888  (fold_0: 38653 seen samples and 7578 unseen samples)\nLocal Score: 0.998~ (old split, single fold)\nPublic LB: 0.9923 (single fold)\n</code></p>\n\n<p>The score got around <code>0.01</code> dropped when there are 16.4% unseen graphemes in the validation set.\nIf anyone is interested in challenging validation with unseen graphemes, please have a look at this post ( <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134434\">new validation setting</a> )</p>",
      "votes": 8,
      "replies": [
        {
          "id": 766734,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-03-08T16:10:47.823000",
          "content": "<p>Thanks for reporting. Public LB is single model, single fold on new validation setting?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 766735,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2020-03-08T16:11:04.083000",
          "content": "<p>water post?😏 😏 <a href=\"/haqishen\">@haqishen</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 766758,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-08T16:43:39.137000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> \nNo, LB score is not like that, before submitting I have to re-train my models on old validation splits (with all grapheme as seen samples) after experimenting on new validating set.\nThen with fully ensemble I got 0.9933.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 766939,
          "author_name": "Gold Retriever",
          "author_url": "",
          "post_date": "2020-03-09T00:41:53.097000",
          "content": "<p>The king is back.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767000,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-09T02:41:43.680000",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> I don't get it. So the seen-unseen split is not a ready-replace plugin for e.g. Random Split? After training on the split we have to train again on old splits? I guess this is for the models to see as many graphemes as possible? Doesn't this introduce some leak into the finetuning stage, as the validation samples in stage-2 might occurred in training samples in stage 1. Wondering how you re-train on old validation splits🤥 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 767010,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-09T03:09:03.473000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Sorry for the confusion, I'll try to make things clear.</p>\n\n<ul>\n<li>There are some unseen graphemes in the public test set, that's why we always have a gap between our CV and LB.</li>\n<li>The gap also indicated that, our models are not good at predicting unseen graphemes.</li>\n<li>We don't know how many unseen graphemes in the private test set, so there is a risk of shake-up.</li>\n<li>To address this problem, a good way is to build a validation split that similar to the public/private test set. That's why we'd better to do validation on [seen + unseen graphemes] set.</li>\n<li>Although I got some progress there, but finally I have no idea how to train my models to generalize to unseen graphemes as well as seen graphemes. </li>\n<li>So although it's kind of pain, I have to re-train all my models on old splits, to make them 'see' as much graphemes as I can.</li>\n</ul>\n\n<p><a href=\"/tonychenxyz\">@tonychenxyz</a> There's no king in Kaggle, we are all learning from each others. It's most important things to us.</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 767016,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-03-09T03:25:28.883000",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> thanks for sharing such insight to tackle the overfitting issue on the private test set.  There are some diverse combinations in Bengali handwriting, a good amount of combination of those three targets and our writing styles. \nHowever, I was thinking to add more <code>grapheme_root</code> samples from the external data sources but I'm not sure whether it will help it or not also it's pain. <strong>BanglaLekha</strong> and <strong>Ekush</strong> data set is pretty good, I think. There are around 98K <code>grapheme_root</code> samples (50 categories, except ঁ ) in <strong>BanglaLekha</strong>. \nWhat do you think? And, I think <a href=\"/tonychenxyz\">@tonychenxyz</a> just wanted to appreciate your success in this competition. 😅 😃 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 767017,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-09T03:26:56.403000",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> Thank you very much for the reply. Here I am observing some anomalies between CV and LB. It seems deeper models perform better on CV but degenerates on LB.</p>\n\n<p>I wonder how you are doing the retraining? How are you able to use stable validation set for retraining?\nI have an idea...We split the unseen graphemes into folds, and train 5-fold</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 767022,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-09T03:43:54.597000",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> I'm not sure how they are organizing the label, and I don't have knowledge about Bengali as well... So I don't know how to do it.</p>\n\n<p><a href=\"/roguekk007\">@roguekk007</a> re-train here I mean training from  imagenet weight again.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 767030,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-09T04:05:33.207000",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> OK so the .9923 LB score has absolutely nothing to do with the new validation scheme...? You train from imagenet weight on the old validation scheme (presumably iterative stratification)? Sorry I was confused. So you are sharing the CV of the unseen validation scheme, which is .9888. I wonder if you have done any experiments submitting this to the LB?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767036,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-09T04:22:57.540000",
          "content": "<p>If you add a trick, and get a same CV score on old split, how do you know it's useful or not?\nnew split can tell you that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 767044,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-09T04:40:30.280000",
          "content": "<p>another way to test is to use some bengali dictionary and write some unseen graphemes by hand and test (even for train)</p>\n\n<p>there are 1296 (not 1295) garphemes here:\n<a href=\"https://github.com/BengaliAI/graphemePrepare/tree/master/collection/A4\">https://github.com/BengaliAI/graphemePrepare/tree/master/collection/A4</a>\n<a href=\"https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt\">https://github.com/BengaliAI/graphemePrepare/blob/master/data/groundTruth.txt</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767075,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-09T06:04:15.943000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  It's hand written recognition competition but not hand writing competition 😂 </p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 767284,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-03-09T12:41:11.530000",
          "content": "<p>I can't read my own writing in French sometimes, I will not risk polluting my model with Bengali hand writing ;)</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 743999,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "2020-02-12T13:16:05.867000",
      "content": "<p>model:resnet34 <br>\nimg_size: 3x 224 x 224(just resize)\naugmentation: cutmix + some augmentations\nCV: 0.985(only first fold, 80% train iterative stratified)\nLB: ???  </p>",
      "votes": 7,
      "replies": [
        {
          "id": 750690,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-19T15:58:06.393000",
          "content": "<p>May i ask what lb score did you get with that model? I am also using resnet34, so i'm curious :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754260,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2020-02-23T09:52:58.703000",
          "content": "<p>Sorry for the late reply. I have never submitted that. <br>\nBecause it's only experimenting for techniques  </p>\n\n<p>Now I just reached 0.988 using ResNet34</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755162,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-24T14:06:32.340000",
          "content": "<p>Great results, I'm doing pretty much exactly like you however can't get to 98%, do you think 3 channels would increase my results? thanks for your advice</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755170,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2020-02-24T14:12:17.983000",
          "content": "<blockquote>\n  <p>do you think 3 channels would increase my results?</p>\n</blockquote>\n\n<p>No. I use 3 channels, but the same values among the image. <br>\npreprocess is grayscale, and I copy the same data to 3 channels.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755246,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-24T15:39:43.727000",
          "content": "<p>Thanks for your advice! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 731260,
      "author_name": "MachineLP",
      "author_url": "",
      "post_date": "2020-01-28T13:14:51.337000",
      "content": "<p>model: se-resnext50 with mixup + cutmix\nsplit: random 80/20\ninput: 3x128x128\nCV : 0.983\nLB:  0.974</p>",
      "votes": 8,
      "replies": [
        {
          "id": 731448,
          "author_name": "Satwik",
          "author_url": "",
          "post_date": "2020-01-28T16:47:36.350000",
          "content": "<p>Hey, what probabilities did you use for mixup and cutmix , and how many epochs did you train for? Also how did you make the image to have 3 channels?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 731733,
          "author_name": "MachineLP",
          "author_url": "",
          "post_date": "2020-01-29T01:14:56.123000",
          "content": "<p>（1）I trained it for 100 epochs。\n（2）img = cv2.cvtColor(img,cv2.COLOR_GRAY2BGR)。</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 733479,
          "author_name": "Youhan Lee",
          "author_url": "",
          "post_date": "2020-01-31T07:19:27.887000",
          "content": "<p>Hi, did you just resize the images(137, 236) to images(128, 128)? or use something special cropping methods?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 770621,
      "author_name": "Taemyung Heo",
      "author_url": "",
      "post_date": "2020-03-13T06:58:50.963000",
      "content": "<blockquote>\n  <p>Model: se_resnext50_32x4d\n  Image Size: 137X236\n  Train Split: 0.8, Single Fold\n  Epoch: 150\n  CV: 0.9976\n  LB: 0.9889</p>\n</blockquote>",
      "votes": 5,
      "replies": [
        {
          "id": 770676,
          "author_name": "ratan rohith",
          "author_url": "",
          "post_date": "2020-03-13T08:45:14.527000",
          "content": "<p>Great <a href=\"/tmheo74\">@tmheo74</a>! Can i know which augmentations are you using? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 770680,
          "author_name": "Taemyung Heo",
          "author_url": "",
          "post_date": "2020-03-13T08:53:46.757000",
          "content": "<p>Yes, we are using cutmix, mixup, gridmask, normal aug(scale, rotate, etc.)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 770694,
          "author_name": "ratan rohith",
          "author_url": "",
          "post_date": "2020-03-13T09:25:50.970000",
          "content": "<p>Thank you <a href=\"/tmheo74\">@tmheo74</a> All the best :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 770724,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2020-03-13T10:17:59.193000",
          "content": "<p>Man, what kind of hardware do you have? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 758464,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-02-27T19:23:55.327000",
      "content": "<p><code>\nmodel: seresnext50 pretrained\naug: ssr + coarse dropout w/ mixup+cutmix\ninput: 3x137x236\nsplit: iterative stratification\ntraining: adam, reducelronplateau, 23 epochs\nsingle fold CV: .97896\nLB: .9711\n</code>\nThe training isn't even done and I have a separate model with CV .985+ I haven't submitted yet as I'm waiting for it to finish training.\nIs there such thing as over sharing? Recently the big tricks have been shared and I think the LB will quickly saturate.\nEDIT:\n.985 model finished here are the results:\n<code>\nsingle fold CV: .99185\nLB: .981 (1.08 gap)\n100 epochs\nSame setup as above model but with a few changes ;)\n</code></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 706203,
      "author_name": "NguyenThanhNhan",
      "author_url": "",
      "post_date": "2019-12-30T04:11:53.830000",
      "content": "<p>model: B0(pretrained=True)\nimg: 1 channel \nimg_sz: original\nsplit: stratified (80/20)\noptim: AdamW\nepoch: 30\nsched: Cosine</p>\n\n<p>CV : 0.9684\nLB:  0.9629</p>\n\n<p>With similar settings, my ResNet50 also got cv 0.9633 and lb 0.9574. \nFocal loss was slightly worse than cross-entropy.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 706273,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2019-12-30T07:01:18.547000",
          "content": "<p>How did you do the stratified split? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 707619,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-01T07:05:30.043000",
          "content": "<p><a href=\"https://www.kaggle.com/yiheng/iterative-stratification\">This</a> might help you.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 761938,
      "author_name": "Zhanseri Ikram",
      "author_url": "",
      "post_date": "2020-03-03T03:50:29.007000",
      "content": "<p>size: 128x128\ncv: 0.9972\nlb: 0.9890\nsingle fold, no tta</p>",
      "votes": 6,
      "replies": [
        {
          "id": 762199,
          "author_name": "Karol Zak",
          "author_url": "",
          "post_date": "2020-03-03T10:15:33.677000",
          "content": "<p>Impressive for 128x128!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 762434,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-03-03T14:04:07.587000",
          "content": "<p><a href=\"/nuller\">@nuller</a> hi, have you used the pre-processed image from <a href=\"https://www.kaggle.com/iafoss/grapheme-imgs-128x128\">here</a>. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762652,
          "author_name": "Zhanseri Ikram",
          "author_url": "",
          "post_date": "2020-03-03T17:25:09.863000",
          "content": "<p>hi <a href=\"/ipythonx\">@ipythonx</a>, no I didn't use his preproc</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 732024,
      "author_name": "Balaji Selvaraj",
      "author_url": "",
      "post_date": "2020-01-29T11:48:55.417000",
      "content": "<p>EfficientNet Models: [Single model , Single Fold]</p>\n\n<p>Things I tried:\nB3 - 300 x 300                                           LB: 0.9639             [40 epochs]\nB3 - 300 x 300    + cutmix                         LB: 0.9695             [40 epochs]</p>\n\n<p>B7 - 128                                                        LB : 0.9683             [40 epochs]\nB7 - 224                                                        LB : 0.9685             [15 epochs]</p>\n\n<p>Things from smaller studies:\n1.Cutmix performs better than mixup for efficientnet series\n2.Still in process of replacing Adaptiveavgpool2d with GeM\n3.Tried out multiple types of tails. Best one was similar to <a href=\"/iafoss\">@iafoss</a> \n4.Tried fade, sprinkles, cutout. Not much improvement.\n5.Using swish inplace of Mish gives me more free space in GPU. Helping me to increase the batch size</p>\n\n<p>Training longer would improve the scores. Currently planning to finalize the model. Then in later stages, to finalize parameters. Then going to run for longer epochs with multiple folds. I hope the approach I mentioned seems decent.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 733358,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-01-31T02:30:38.850000",
          "content": "<p>Hi! I am using efficientnet too. I wonder what LR and batch size you are using? Also, from my experiments cutmix and mixup perform best together (.5 .5 probability), maybe you should try that. Also, I wonder if you have any success implementing GeM?? For me, GeM results in training error after several epochs (floatpoint div by 0 error when using apex). I recommend you to check out the github mishcuda repository, I saw marginal improvement with mish.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 733397,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-31T04:23:20.277000",
          "content": "<p>Hey! I tried efficientnet but it seems like i can't get a good score compared to resnets. I cap at .95-.96 validation. I added 3 tails with conv/batchnorm just like my resnet models, i tried with GeM too but doesnt seem to work.. Do you have any advice?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 733428,
          "author_name": "Balaji Selvaraj",
          "author_url": "",
          "post_date": "2020-01-31T05:28:31.090000",
          "content": "<p>I remodified the GeM and its working better than adpativeavgpool2d as mentioned here. I have tested for only one epoch runs. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 733579,
          "author_name": "Balaji Selvaraj",
          "author_url": "",
          "post_date": "2020-01-31T10:17:45.220000",
          "content": "<p>Are you using rwightman implementation or lukemelas implementation</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 734213,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-02-01T05:07:25.497000",
          "content": "<p>Im using this implementation of GeM:</p>\n\n<p>```\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def <strong>init</strong>(self, p=3, eps=1e-6):\n        super(GeM,self).<strong>init</strong>()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def <strong>repr</strong>(self):\n        return self.<strong>class</strong>._<em>name</em>_ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'</p>\n\n<p>```\nWhich is working well on other models. Have you tried a cnn/batchnorm tail with efficientnet? I tried but it doesnt seem to do so well</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 734254,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-02-01T07:01:22.557000",
          "content": "<p>Also efficient, b4 with 224x224, rotate, cutmix and mixup, finally lb 0.9682 (100 epochs) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 736441,
          "author_name": "Balaji Selvaraj",
          "author_url": "",
          "post_date": "2020-02-04T07:22:52.693000",
          "content": "<p>Most models will get WxH in the range of 16x16 or something similar. So if you use \"Conv\" in the tails it will work. But for efficientnet. Its 2560x4x4.  So using a conv layer on the tail doesn't going to capture much information.</p>\n\n<p>So dont prefer conv layer for efficientnet. I am opening a thread specific for efficientnet ppl. So we can discuss and improve the score. </p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128911\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128911</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 726402,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2020-01-23T01:56:50.750000",
      "content": "<p>model: xresnet34(pretrained=False)\nimage size: 128\noptim: Lookahead(SGD)\nsched: OneCycleLr\naugmentations: RandomResizedCrop, RandomRotate, RandomErasing, RandomPerspective</p>\n\n<p>CV: 0.9740\nLB: 0.9672</p>",
      "votes": 5,
      "replies": [
        {
          "id": 730135,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-01-27T06:20:23.127000",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Wow! That is a very charming score considering your configuration. I was not able to reach that score without using some magic. Do you mind telling if you used some tricks for this score?</p>\n\n<p>Also, my LB and CV correlation is nearly identical as yours: ~.007 gap. Cheers</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 730461,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-27T14:08:26.393000",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> What i found worked the best after many tests:\n -3 tails, one for each category\n-Adding a weight tensor to the loss function to handle imbalanced classes\n-Training for longer (60+ epochs)\n-Training with SGD and OneCycleLr gave me better results with the same setup</p>\n\n<p>Other than that just hyperparameter tuning!</p>\n\n<p>I seem to not be able to have a better result with mixup/cutmix, do you have any advice? What magic did you use? :p</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 730823,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-01-28T01:20:07.213000",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> \nI am using cutmix and mixup with .5 probability each :) Here is my implementation. There highly-voted posts out there, too. For me, I found a .0005 moderate improvement. With onecycle I also observe better scores, but I have to tune the number of epochs (which is extremely time-consuming). Still testing out how much more training cutmix+mixup requires. For me, the rule of the thumb is 1.5-2 times more training.</p>\n\n<p>Copied and edited from a wide range of sources\n```\ndef rand_bbox(size, lam):\n    # Revision because there is only 1 channel\n    W = size[1]\n    H = size[2]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)</p>\n\n<pre><code># uniform\ncx = np.random.randint(W)\ncy = np.random.randint(H)\n\nbbx1 = np.clip(cx - cut_w // 2, 0, W)\nbby1 = np.clip(cy - cut_h // 2, 0, H)\nbbx2 = np.clip(cx + cut_w // 2, 0, W)\nbby2 = np.clip(cy + cut_h // 2, 0, H)\n\nreturn bbx1, bby1, bbx2, bby2\n</code></pre>\n\n<p><code>\n</code></p>\n\n<h1>For training loop</h1>\n\n<pre><code>    x, labels1, labels2, labels3 = batch\n   # Cutmixup here\n    choice = np.random.rand(1)\n    if choice &amp;lt;= args['cutmix']:\n        # Cutmix!\n        lam = np.random.beta(1., 1.)\n        rand_index = torch.randperm(x.shape[0]).cuda()\n        l1_a, l2_a, l3_a = labels1, labels2, labels3\n        l1_b, l2_b, l3_b = l1_a[rand_index], l2_a[rand_index], l3_a[rand_index]\n        bbx1, bby1, bbx2, bby2 = rand_bbox(x.size(), lam)\n        x[:, bbx1:bbx2, bby1:bby2] = x[rand_index, bbx1:bbx2, bby1:bby2]\n        # adjust lambda to exactly match pixel ratio\n        lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (x.size()[-1] * x.size()[-2]))\n        # compute output\n        preds1, preds2, preds3 = model(x)\n        loss = criterion(preds1, preds2, preds3, l1_a, l2_a, l3_a) * lam + \\\n                criterion(preds1, preds2, preds3, l1_b, l2_b, l3_b) * (1 - lam)\n    elif choice &amp;lt;= args['cutmix'] + args['mixup']:\n        # Mixup!\n        indices = torch.randperm(x.size(0))\n        shuffled_x = x[indices]\n        shuffled_l1, shuffled_l2, shuffled_l3 = labels1[indices], labels2[indices], labels3[indices]\n        lam = np.random.beta(.4, .4)\n        x = x * lam + shuffled_x * (1 - lam)\n        preds1, preds2, preds3 = model(x)\n        loss = lam * criterion(preds1, preds2, preds3, labels1, labels2, labels3) +\\\n                        (1 - lam) * criterion(preds1, preds2, preds3, shuffled_l1, shuffled_l2, shuffled_l3)\n    else:\n        preds1, preds2, preds3 = model(x)\n        loss = criterion(preds1, preds2, preds3, labels1, labels2, labels3)\n\n    with amp.scale_loss(loss, op) as scaled_loss:\n        scaled_loss.backward()\n</code></pre>\n\n<p>```</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 730833,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-28T01:52:56.803000",
          "content": "<p>Thanks! I saw that post, and tried it but it seems like i couldn't get a better score, maybe i didnt train long enough.. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 715130,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2020-01-10T06:40:17.173000",
      "content": "<p><code>\nmodel: se-resnext50\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.977519\nLB:  0.9682\n</code></p>",
      "votes": 5,
      "replies": [
        {
          "id": 716017,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-11T06:36:16.717000",
          "content": "<p>So you didn't resize the image rather than go with raw size! Have you used any image preprocessing of your own? or data augmentation? Seeing your CV score, how many folds did you make? If you made, let's say 5 fold and got 5 models, then did you pick up the single best scoring model for submission? </p>\n\n<p>However, CV implementation is pesky for the multi-label problem. Would you like to please view this thread where I've mentioned a <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#715487\">CV issue</a>? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716021,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-01-11T06:48:10.987000",
          "content": "<p>I didn't use any image preprocessing. this is a single fold score based on 80/20 split of data. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716187,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-11T11:27:36.323000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 763286,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T10:24:36.987000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 719121,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-15T06:45:32.930000",
      "content": "<p><code>\nif np.random.rand()&amp;lt;0.5: do mixup\nelse: do cutmix\n</code></p>\n\n<p>i am thinking of</p>\n\n<p>```\nrandom.choice([mixup, cutmix, cutout, sprinkle, augmix, etc ...])</p>\n\n<p>```</p>",
      "votes": 6,
      "replies": [
        {
          "id": 721603,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-17T14:22:53.327000",
          "content": "<p>cutout helps. </p>\n\n<p><code>\nmodel: se-resnext50 (pertained= True)\naugment: cutout, rotate(20)\nsplit: random 80/20\ninput: 1 x 128 x 128\nValidation Score : 0.9821900\nLB:  0.9745\n</code>\nwithout cutout ruining with the same set up I was getting LB <code>0.969</code></p>\n\n<p>Image preprocessing - with <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a></p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 721643,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-17T15:01:26.963000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721644,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-17T15:02:02.857000",
          "content": "<p>Impressive. However, One thing I like clarify. As you said you split the training set 80:20, how you state validation score as a CV score. Isn’t it when we do KFold cross validation, then we say about CV score! So, is it cross validation (CV) or simply just the validation score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721646,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-17T15:04:25.883000",
          "content": "<p>Ahhh... Thanks for bringing this up. I guess I should be clear. Its a validation score not 5 fold CV =) I will correct </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721657,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-17T15:22:59.687000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721715,
          "author_name": "Mahtab Noor Shaan",
          "author_url": "",
          "post_date": "2020-01-17T16:34:04.257000",
          "content": "<p>I cannot use image size of 128 x 128 with pretrained se-resnext models due to problems with somewhere the image is resized into say 2048 x -2 x -2 :p I found that the existing architecture indeed results in this error. So, can you give me any advice on how you do this? <a href=\"https://www.kaggle.com/drhabib\"></a><a href=\"/drhabib\">@drhabib</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 721724,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-17T16:42:00.323000",
          "content": "<p>Ahh  are you using Pytorch ? </p>\n\n<p>it seems like you have to add Polling Layer and than flatten. </p>\n\n<p><code>\n1) First cut the model\n2) add  nn.AdaptiveAvgPool2d(1) \n3) Flatten ()\n4) nn.Linear\n</code></p>\n\n<p>Let me know if you need more detailed example</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 721771,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-17T17:18:33.930000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> thanks for your clarification. However, one thing more, the validation score; well I'm assuming you've used softmax and cross_entropy_loss. So, you should have three validation score, such as: grapheme root, vowel, and consonant. So, do you mention the average validation score of them? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721779,
          "author_name": "Mahtab Noor Shaan",
          "author_url": "",
          "post_date": "2020-01-17T17:26:16.583000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Thank you very much for the explanation! I think I got it. I might bug you again if I face any problem! :p Sorry in advance.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721788,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-17T17:38:06.220000",
          "content": "<p><a href=\"/ipythonx\">@ipythonx</a> \nOk so I just train one model with 80% being training data and 20% validation.</p>\n\n<p>scores for individual groups:</p>\n\n<p><code>\ngrapheme - 0.976273\nvowel - 0.990575\nconstant  - 0.985638\n</code></p>\n\n<p><code>\nlocal score using competition metric: 0.9821900\nlb scores: 0.9745\n</code></p>\n\n<p>I hope its clear =) </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 721794,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-17T17:44:08.530000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> That sure does. Thank you. 😃 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721850,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-01-17T19:30:32.470000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> thanks, I have a few questions : what alpha do you use for cutout (for the beta distribution)?  do you perform it at every batch (or only x% of batches)? what does your training score looks like (can you still overfit)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721880,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-17T20:10:02.503000",
          "content": "<p>actually one can google for \"mixup\" and\"cutout\" for more advanced versions or other variations. find papers that cite the original papers.</p>\n\n<p>e.g. here is a variation just published</p>\n\n<p>GridMask Data Augmentation\nP Chen - arXiv preprint arXiv:2001.04086, 2020 - arxiv.org\n3 days ago - We propose a novel data augmentation methodGridMask'in this paper. It\nutilizes information removal to achieve state-of-the-art results in a variety of computer vision\ntasks. We analyze the requirement of information dropping. Then we show limitation of …</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 722002,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-18T01:22:16.773000",
          "content": "<p>Thanks <a href=\"/hengck23\">@hengck23</a>  . And for everyones use here is the official code for Gridmask </p>\n\n<p><a href=\"https://github.com/akuxcw/GridMask\">https://github.com/akuxcw/GridMask</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 722698,
          "author_name": "Ee Kin Chin",
          "author_url": "",
          "post_date": "2020-01-18T23:58:49.310000",
          "content": "<p>here is some random gridmask applied to the dataset <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1203323%2F08b26252532a94ecf5028b025d66bb19%2Fgridmask.png?generation=1579391912490101&amp;alt=media\" alt=\"\"></p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 723906,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2020-01-20T15:37:56.140000",
          "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a> , it seems that I am severely overfitting. Can u tell what optimizer, scheduler u used? What are ur learning rates and any weight decay ? Thanks. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 723943,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-20T16:07:46.173000",
          "content": "<p>Hi <a href=\"/virajbagal\">@virajbagal</a> <br>\nI would say you should worry very little about optimizers, schedulers and learning rates. For example I can tell you can reach top 30 with Adam or SDG with learning rate anywhere in between (0.01-0.003). </p>\n\n<p>I would also recommend to choose one model, e.g Resnet50 or Seresnext (more bigger) and do all your experiments. </p>\n\n<p>You correctly identified the problem which is overfitting on training data. How this problem can be solved ? I recommend you go thru all the highly upvoted posts in this thread and write down there approaches. Make a list and go one by one. There is no other way.  </p>\n\n<p>After you done this, you can focus on optimizers,  hyperparamters and bigger models. </p>\n\n<p>Hope to see you in top 50 =) </p>",
          "votes": 19,
          "replies": []
        },
        {
          "id": 723952,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-20T16:27:26.077000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a>  i should probably print your golden suggestion .</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 723956,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-20T16:35:48.527000",
          "content": "<p>You should be giving advice to us, since you are in top 10 =) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 724544,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2020-01-21T08:28:04.873000",
          "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a> , thank you for ur advice. I am following the top posts but I am not able to reproduce the results. It seems many of u guys have easily reached 0.96+ using only simple augs, random splits and normal learning methods,  but here I am struggling to even cross 0.9595 . I don't understand where I am going wrong. Anyways, it is all about trying out I guess. I'll try even harder xD </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 724791,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-21T13:42:37.170000",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> I see the frustration =)  I would recommend then looking on top 3 scoring kernels with 0.96+ and try to tweak them. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 726441,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-01-23T02:26:25.053000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> when you use a pretrained model do you freeze some layers and then train? Thanks in advance!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 727195,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-23T14:21:19.360000",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> it did not help so much. so I <code>unfreeze</code> everything and train.</p>\n\n<p>Good luck</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 703378,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-26T04:22:01.640000",
      "content": "",
      "votes": 5,
      "replies": [
        {
          "id": 704683,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-27T20:38:56.207000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 706222,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-30T05:07:48.973000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 762152,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-03T08:56:44.057000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 763435,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T13:41:00.803000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 765396,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-06T15:24:31.833000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 756846,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-26T06:23:07.660000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 757456,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-26T19:28:34.260000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 758157,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-27T13:50:07.133000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 732272,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-29T16:23:48.797000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 711915,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-06T16:50:00.230000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 712856,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-07T16:35:37.590000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 712903,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-07T17:47:24.017000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 712913,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-07T18:15:45.243000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 713294,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-08T05:49:50.710000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 713359,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-08T07:40:14.990000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 713620,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-08T13:32:04.197000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714338,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-09T10:17:51.990000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714748,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-09T17:57:17.670000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714752,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-09T18:02:53.480000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 714756,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-09T18:06:59.657000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 715687,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-10T18:14:25.657000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715775,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-10T20:01:56.677000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715798,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-10T21:12:11.863000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716644,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-12T03:46:30.130000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 716838,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-12T10:49:32.457000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 716840,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-12T10:51:03.440000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 717584,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-13T10:27:07.393000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 719049,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-15T04:16:33.460000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 723859,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-20T14:43:06.043000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 734858,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-02T05:22:49.393000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 735331,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-02T22:12:45.803000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 705279,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-28T17:46:42.823000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 705410,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-28T22:40:47.857000",
          "content": "",
          "votes": 5,
          "replies": []
        },
        {
          "id": 705818,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-29T14:20:25.077000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 705823,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-29T14:29:59.477000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 705826,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-29T14:40:48.417000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 705828,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-29T14:48:10.173000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 738999,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-07T09:15:37.720000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 704990,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-28T08:59:25.567000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 706105,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-29T23:27:43.887000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 703421,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-26T06:16:57.530000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 710751,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-05T07:03:47.857000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 710758,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-05T07:17:10.543000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 707575,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-01T04:58:20.913000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 707794,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-01T13:58:52.457000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 704943,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-28T07:30:34.463000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 707207,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-31T12:17:24.137000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 703276,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-25T23:12:22.040000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 703414,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-26T06:12:51.777000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 703420,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-26T06:16:02.540000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 703423,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-26T06:18:41.503000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 703425,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-26T06:20:38.160000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 705460,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-29T00:09:26.290000",
          "content": "",
          "votes": 19,
          "replies": []
        }
      ]
    },
    {
      "id": 759655,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-29T09:35:24.437000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 759898,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-29T15:40:39.457000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759976,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-29T17:21:41.200000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 763042,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T04:13:54.303000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 741811,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-11T00:42:50.960000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 768182,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-10T14:17:22.283000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 768247,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-10T15:26:01.873000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 768411,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-10T19:04:23.537000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 769188,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-11T16:12:34.560000",
          "content": "",
          "votes": -1,
          "replies": []
        },
        {
          "id": 769235,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-11T17:06:36",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 761813,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-03T00:29:54.730000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 768222,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-10T14:58:19.763000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 768330,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-10T16:42:55.663000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 761792,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-03T00:00:39.750000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 764209,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-05T08:17:21.170000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764227,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-05T08:28:52.140000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 768333,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-10T16:46:44.437000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 725802,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-22T14:06:59.213000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 722046,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-18T03:11:52.317000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 711664,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-06T11:19:07.270000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 717443,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-13T05:43:43.117000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717454,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-13T06:15:23.807000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 708473,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-02T10:29:49.577000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 708487,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-02T10:50:28.597000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 708574,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-02T12:51:21.160000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 709230,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-03T07:57:13.150000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 709267,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-03T09:50:13.530000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 709312,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-03T10:52:51.317000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 770620,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-13T06:58:28.727000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 768540,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-10T23:54:07.830000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 763422,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-04T13:22:45.220000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 763427,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T13:33:20.733000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 763460,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T13:59:38.673000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 763466,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T14:05:37.740000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 763783,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T21:08:08.787000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763798,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T21:51:02.837000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764396,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-05T12:21:09.853000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764483,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-05T14:05:00.573000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764552,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-05T15:34:26.183000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 764808,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-05T23:31:57.757000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 761109,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-02T06:37:36.253000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 761114,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-02T06:48:37.687000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 761127,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-02T07:07:05.600000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 761845,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-03T01:33:00.387000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 757530,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-26T21:53:16.410000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 759003,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-28T12:48:42.113000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 759016,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-28T13:03:43.307000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763291,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T10:27:38.760000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 724404,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-21T06:07:49.580000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 724444,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-21T06:45:55.470000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 730314,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-27T10:42:39.937000",
          "content": "",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 710682,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-05T04:38:50.803000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1006617,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-11T12:21:21.710000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767900,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-10T08:17:02.373000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 768211,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-10T14:46:52.207000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 769753,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-12T08:19:55.817000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 769756,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-12T08:20:54.423000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 771247,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-13T23:27:19.523000",
          "content": "",
          "votes": -2,
          "replies": []
        }
      ]
    },
    {
      "id": 762046,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-03T06:44:29.980000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 753070,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-21T17:19:43.697000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 753127,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-21T18:38:03.577000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 761066,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-02T04:59:15.577000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 759957,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-29T16:49:44.367000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 742605,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-11T12:07:23.333000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 704188,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-27T06:32:19.100000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 710146,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-04T11:05:38.943000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "703111": "Just starting a common competition thread =) \nWhat is your current best single model?\nMy:\n\n```\nmodel: xresnet18(pretrained=False)\nimg: 1 channel \nimg_sz: 128\nsplit: random (80/20)\noptim: Adam\nepoch: 15\nsched: Cosine decay\n\nCV : 0.9630\nLB:  0.9611\n```",
    "718926": "[Update]\n```\nmodel: se-resnext50 with mixup + cutmix\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.988011\nLB:  0.9790\n```\n```python\nif np.random.rand()&lt;0.5: do mixup\nelse: do cutmix\n```",
    "717760": "The submission of\n`LB 0.9875`\nwas a single fold model which scored on my local experiment by\n`CV 0.996792`\n\n\nHope this will be motivation for you guys to improve the performance of single model, good luck!",
    "723165": "CV: 0.9913\nLB: 0.9848\n\nthough I still haven't reached a single-model score as high as @haqishen but I haven't used any TTA and big boys like seresnext or high efficientnet/densenet due to free time and computational difficulties so I think 0.99 from single model is very achievable. Have fun single model racing guys :D",
    "742587": "model: se-resnext50\nimg_size: 3x137x236\naugmentation: rotate, cutmix \nCV : 0.994\nLB:  0.985\n\nI can not get the CV score more than 0.997 from some top kagglers. So perhaps I can get some advice and help from your guys here😃 ",
    "734490": "I met kernel error, so I couldn't submit :(\n```\n- img_size: 137x236\n- augmentation: auto augment, augmix\n- OHEM\n\ncv: 0.997\nLB: 0.9846\n```",
    "756554": "Model: Densenet121\nimg_size: 3x224x224 (simple resize)\nAugmentation: NOT cutmix or mixup\nEpoch: 40 (still running) \nCV: 0.9938\nLB: 0.9825\n\nTo my surprise, not all augmentation methods suit all architectures. You might have to find out which augmentation will work best for your model.",
    "733514": "cv: 0.9890\nlb: 0.9794\n\n```\nmodel: se_resnext50_32x4d\nimgsize: 128x128\nsplit: 5/6 train, 1/6 valid\ninference: 15 minutes (kaggle kernels)\nno tta, no ensemble\n```\n",
    "762424": "`update`\ncv: 0.9981\nlb: 0.9900\nsingle fold, no tta\nI still haven't found [haqishen‘s](https://www.kaggle.com/haqishen) magic😑 😑 ",
    "735506": "local: `0.9970`\nlb: `0.9884`\n\nsingle model, single fold, no tta\n\ni still think `0.99` is really doable but the you have to use big models/ tta and maybe some lucky with random seed because improvement seems quite random when local error is small",
    "761883": "You guys may believe or not. \nYesterday, I dreamed about a solution of the top team. Even I could not remember the details, however, when I waked up, I had an idea to follow up. It gave me a major improvement 😂  \n\nimg_size: 137x236\nCV: 0.993 \nLB: 0.9868 \nSingle fold, no tta. ",
    "738975": "Model: se-resnext50-32x4d\nImage size: 128x128x1\nCV: 0.9937\nLB: 0.9838\n\nInteresting that no matter how I changed model structure, augmentation or image size,\nLB scores are always equal to my CV scores minus about 1~1.3%, \nguess I need totally different way to break through 99%",
    "737631": "**CV : 0.979**\n**LB : 0.971**\n```\nModel: SEResNeXT50\nAugmentations: 40% Mixup + 40% Cutmix + 10% Cutout + 10% GridMask\nSplit: Character stratified split 80/20\nImage size: 128x128x3 (just resize)\nEpoch: 100\nSingle fold, No TTA\n```\n\n[Next]\n- AugMix\n- Reduced Focal Loss\n- Bigger image size\n...\n\nThere are lot of things to be done before getting to 0.99...\n\nI would appreciate if you could give me any advice!!",
    "766744": "```\nSingle Fold:\nCV: 0.9967\nLB: 0.9884 \n```\n\nEfficientNetB4, 130x224, cutmix, OHEM ",
    "761824": "img_size: 137x236\ncv: 0.996\nlb: 0.9894\nsingle fold, no tta",
    "772733": "Finally reached cv0.9975, but it's too late...\nmodel: efficientnet-b3\nsplit: 5fold with stratified k-fold\nimage size: 224x224\naugment: gridmask, cutout",
    "721664": "I completed initial analysis using cam class activation map. Indeed there is much overfitting. Training loss can be 0.01 and validation loss can be 0.20 to 0.13. With regularisation like mixup and cutout, the cam is vastly improved.  \n\nHence to get good results, one can rely on such regularization or add hand more label signal ( e.g pixel or box annotations) to improve the cam response",
    "716280": "[Update]\n```\nmodel: se-resnext50 with mixup\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.982066\nLB:  0.9739\n```",
    "703644": "model: ResNet 18 (pretrained=False)\nimg: 1 channel \n\nCV: 0.9654\nLB:  0.9632\n\nI think CV and LB are correlated well, according to some posts in this thread :-)\n\n---\n\nupdated on Jan.02 2020\n\nResNet 18\nCV: 0.9729\nLB: 0.9681",
    "766728": "My score on [new validation setting](https://www.kaggle.com/c/bengaliai-cv19/discussion/134434) is like this:\n\n```\nLocal Score: 0.9888  (fold_0: 38653 seen samples and 7578 unseen samples)\nLocal Score: 0.998~ (old split, single fold)\nPublic LB: 0.9923 (single fold)\n```\n\nThe score got around `0.01` dropped when there are 16.4% unseen graphemes in the validation set.\nIf anyone is interested in challenging validation with unseen graphemes, please have a look at this post ( [new validation setting](https://www.kaggle.com/c/bengaliai-cv19/discussion/134434) )",
    "743999": "model:resnet34  \nimg_size: 3x 224 x 224(just resize)\naugmentation: cutmix + some augmentations\nCV: 0.985(only first fold, 80% train iterative stratified)\nLB: ???  ",
    "731260": "model: se-resnext50 with mixup + cutmix\nsplit: random 80/20\ninput: 3x128x128\nCV : 0.983\nLB:  0.974",
    "770621": "&gt; Model: se\\_resnext50\\_32x4d\nImage Size: 137X236\nTrain Split: 0.8, Single Fold\nEpoch: 150\nCV: 0.9976\nLB: 0.9889",
    "758464": "```\nmodel: seresnext50 pretrained\naug: ssr + coarse dropout w/ mixup+cutmix\ninput: 3x137x236\nsplit: iterative stratification\ntraining: adam, reducelronplateau, 23 epochs\nsingle fold CV: .97896\nLB: .9711\n```\nThe training isn't even done and I have a separate model with CV .985+ I haven't submitted yet as I'm waiting for it to finish training.\nIs there such thing as over sharing? Recently the big tricks have been shared and I think the LB will quickly saturate.\nEDIT:\n.985 model finished here are the results:\n```\nsingle fold CV: .99185\nLB: .981 (1.08 gap)\n100 epochs\nSame setup as above model but with a few changes ;)\n```",
    "706203": "model: B0(pretrained=True)\nimg: 1 channel \nimg_sz: original\nsplit: stratified (80/20)\noptim: AdamW\nepoch: 30\nsched: Cosine\n\nCV : 0.9684\nLB:  0.9629\n\nWith similar settings, my ResNet50 also got cv 0.9633 and lb 0.9574. \nFocal loss was slightly worse than cross-entropy.",
    "761938": "size: 128x128\ncv: 0.9972\nlb: 0.9890\nsingle fold, no tta",
    "732024": "EfficientNet Models: [Single model , Single Fold]\n\nThings I tried:\nB3 - 300 x 300                                           LB: 0.9639             [40 epochs]\nB3 - 300 x 300    + cutmix                         LB: 0.9695             [40 epochs]\n\nB7 - 128                                                        LB : 0.9683             [40 epochs]\nB7 - 224                                                        LB : 0.9685             [15 epochs]\n\nThings from smaller studies:\n1.Cutmix performs better than mixup for efficientnet series\n2.Still in process of replacing Adaptiveavgpool2d with GeM\n3.Tried out multiple types of tails. Best one was similar to @iafoss \n4.Tried fade, sprinkles, cutout. Not much improvement.\n5.Using swish inplace of Mish gives me more free space in GPU. Helping me to increase the batch size\n\nTraining longer would improve the scores. Currently planning to finalize the model. Then in later stages, to finalize parameters. Then going to run for longer epochs with multiple folds. I hope the approach I mentioned seems decent.\n\n",
    "726402": "model: xresnet34(pretrained=False)\nimage size: 128\noptim: Lookahead(SGD)\nsched: OneCycleLr\naugmentations: RandomResizedCrop, RandomRotate, RandomErasing, RandomPerspective\n\nCV: 0.9740\nLB: 0.9672",
    "715130": "```\nmodel: se-resnext50\nsplit: random 80/20\ninput: 3x137x236\nCV : 0.977519\nLB:  0.9682\n```",
    "719121": "```\nif np.random.rand()&lt;0.5: do mixup\nelse: do cutmix\n```\n\ni am thinking of\n\n```\nrandom.choice([mixup, cutmix, cutout, sprinkle, augmix, etc ...])\n\n```",
    "703378": "mine:\n```\nresnet34\nimg_sz: 128\nsched: OneCycleWithWarmup\nsplit: random (80/20)\nimg: 3 channels\nepoch: 15\nCV: 0.9610\nLB: 0.9643\n```",
    "762152": "CV: 0.9890\nLB: 0.9845\n\nLooking for +1 teammember with higher CV and LB",
    "756846": "model: se-resnext50\nimg_size: 128, chenel: 3\naugumentation: cutmix * autoaugment\nepoch:100\ncv: 0.984\nlb: 0.9786\n\nAround next week, I can use gpu(v100).\nSo I want to try img size 224.",
    "732272": "Update 2\n```\nmodel: seresnext50 pretrained w/ mixup+cutmix\naffine augmentation\ninput: 3x128x128\nsplit: iterative stratification\ntraining: adam, reducelronplateau, 50-70 epochs\nCV: .97177\nLB: .9643\n```\nUpdate 3\n```\nsame model\n.4 mixup, .4 cutmix, .2 aug only\naffine aug\nsame input\nsame split\nadamw, one cycle lr, 100 epochs\nCV: .97699\nLB: .9661\n```\nsame tail as https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964",
    "711915": "Update:\n```\nmodel: se-resnext50\nAugmentations: Random rotate(+-15 degree), Zoom(1.1), Cutout etc.\nimg: 1 channel \ndim: 128\nEpoch: 25\noptim: over9000\nCV : 0.9704\nLB:  0.9651\n```",
    "734858": "I would like to add this as comparison to ohem https://github.com/qychen13/DifficultyAwareEmbedding",
    "705279": "@drhabib   Have u used any augmentations ? ",
    "704990": "I published [Bengali: SEResNeXt prediction with pytorch](https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch) which achieves 0.9583.\n\n```\nmodel: seresnext101_32x4d (pretrained=True)\nimg: 1 channel \nimg_sz: 128\nsplit: random (90/10)\noptim: Adam\n\nLB:  0.9583\n```\n\nThis is my first submit and since I have not tuned so much, I think the score can be improved more.",
    "703421": "model: resnet34(pretrained=False)\nimg: 1 channel \nimg_sz: 112\nsplit: random (50/50)- trained only on 168*5 images per class\noptim: Adam\nepoch: 100 +EarlyStopping ( patience = 10)\nsched: -\npreprocesssing - cropping to character + resize to 112x112\nCV : 0.9588\nLB:  0.9373",
    "710751": "I got 0.9671 on public LB with a single DenseNet121.\n\n```\nmodel: DenseNet-121 (modified some parts)\nimg: 1 channel \ndim: 128\nsplit: random 80/20\noptim: SGD\nepoch: 60\n```",
    "707575": "model: resnet34(pretrained=True)\nimg: 3 channel \nimg_sz: 128\n\nCV : 0.9650\nLB:  0.9670",
    "704943": "Hey Folks, \nAm i missing something or data seems to be \"easy\" that everyone has so high metric scores straightaway? \nPlus LB in somewhat sync with CV stealing the fun!\nPS Newbie to this comp..",
    "703276": "How did you make the model accept gray scale images? From keras.applications, they only accept RGB images.",
    "759655": "(update)Model: Seresnet50\nimg_size: 3x137 * 236\nAugmentation:  cutmix + cutout\nCV: 0.997\nLB: 0.989",
    "741811": "Hi @drhabib I was wondering if you could share your code for cv 0.9630 lb 0.9611 model. I wonder how you could make the gap between cv and lb so small. I currently do 80/20 stratified shuffle split and reached 0.9886 cv, but lb is only 0.9679, so I would really appreciate if you could publish you low-score pipeline. Thank you very much!",
    "768182": "model: SE-ResNeXt50 with DropBlock\ninput size: 128x128\nvalidation: iterative stratification (single fold)\n\nCV=0.9888, LB=0.9758\nCV=0.9905, LB=0.9764\n\nLarge CV/LB gap...",
    "761813": "Model: DenseNet121\nSize: 224x224\nAug: CO, GM\nCV: 0.9936\nLB: 0.9842\n\nI still have few ideas that I'm experimenting with right now.\nI might be looking for a teammate with a high scoring seresnext50 model (I didn't have any luck with it).\nPM me if anyone would be interested",
    "761792": "Model: se-resnext50\nimgSize: 128x128\nAug: Cutout\nEpoch: 200\nLocal validation 0.9921\nLB: 0.9815",
    "725802": "CV: 97.03%\nLB: 95.64%\n\n```\nmodel: DeepCNN\nimg: 1 channel \nimg_sz: 64\nsplit: random (80/20)\noptim: Adam\nepoch: 15\nrotate:20\n```\n",
    "722046": "@bibek777 Hi! Thank you for sharing your wonderful results. Just wondering, when doing cutmix / mixup is there probability of using usual training at all? So far as I know, implementation of mixup in fast.ai only uses mixup with .4 probability. With \nif np.random.rand()&lt;0.5: do mixup\nelse: do cutmix\nThe model is purely trained using samples mixed with mixup or cutmix?",
    "711664": "eff0 \ncv5: 0.9689\nlb:  0.9615",
    "708473": "one of my tries\n\n```\nmodel: B0(pretrained=True)\nimg: 3 channel\nimg_sz: 256\nsplit: random split (80/20)\noptim: Over9000\nepoch: 20\nsched: Cosine w/o warmup\n```\n\nI also tried B2 and B4 models but they didn't get a boost. \nMaybe more tuning is needed or just model is just big enough to overfit.\n\nCV : 0.9703\nLB : 0.9640",
    "770620": "augmentation: Cutmix + cutout + grid\ninput size: 137 x 236\nsplit: 97.5% + 2.5%\n\nModel/CV/LB: SE-ResNext50/99.38/98.57 with OHEM\nModel/CV/LB: SE-ResNext50/99.59/98.54 w/o OHEM\nModel/CV/LB: B3/99.20/98.18 with OHEM",
    "768540": "i wondered if anyone has results with models using dilated convolution?",
    "763422": "resnet34: single fold(all train data)\nLB: 0.9807\n\nEfficientNet-b2: iterative stratification(train:valid=8:2), not 5 fold average\nCV: 0.9916\nLB: 0.9865\n\nI'm afraid of overfitting. Have a nice day!",
    "761109": "model:seresnext50\nsize:3x137x236\naugumentation: cutmix etc\nCV:0.9965\nLB:0.9866\nLooking for a teammate!",
    "757530": "Model: EfficientNetB3\nimg_size: 3x102x177 (just scaled with factor 0.75)\nAugmentation: No Augmentation\nEpoch: 60\nCV: 0.977 (using validation score for root recall..I noticed this gives a way better indication then the validation score for the official metric.\nLB: 0.9724",
    "724404": "i wonder is there any comparison with and without using imagenet pretrained models?\ni have been using pretrained models for my work and is curious about the performance for training from scratch",
    "710682": "```\nmodel: resnet18\nimg: 1 channel \ndim: 128\nsplit: iterative stratified split(80/20)\noptim: over9000\nepoch: 32\nCV : 0.9676\nLB:  0.9604\n```\nThe gap between my CV and LB is always around 0.007. Any suggestion regarding train-val split to get similar scores would be appreciated. ",
    "1006617": "nb............",
    "767900": "CV 0.988 LB ?\nI'm new to this competition, still cannot submit Lol",
    "762046": "CV : 0.9775\nLB:  0.9747\nseresnext50\nStill trying to use different augs, looking forward to improve score more",
    "753070": "Is there any notebook that uses cutmix in PyTorch. There are no such notebooks in the kernel section. Can anyone share?",
    "761066": "",
    "759957": "",
    "742605": "",
    "704188": "",
    "710146": "Thanks for sharing... "
  }
}