{
  "id": 135959,
  "title": "Overview of my solution",
  "url": "/competitions/bengaliai-cv19/discussion/135959",
  "author_name": "",
  "post_date": "2020-03-16T23:40:01.750519700Z",
  "votes": 24,
  "comment_count": 17,
  "views": 0,
  "content": "<p>There are no insights into my solution that could ruin the LB.  I think I can share now.</p>\n\n<p>Overview of my solution:\n- 5folds, random split (seems the main problem)\n- Original size - 236,137, Also tried cropping, resizing\n- Apex, O0 opt level (pure fp32). For some reason O1 opt level worsened score\n- Best models: efficientnet b5 noisy student, wsl_resnext101. It’s my first time then efficientnet works! \n- AdamW, initLR 3e-4, minLR 1e-6\n- Warmup 3 epochs + Cosine Decay 200 epochs + Cooldown 10 epochs. Didn’t play much with scheduling.\n- Mixup + Cutmix, Cutmix-prob 1.0, Cutmix-alpha 1.0, Mixup-prob 0.4, Mixup-alpha 0.2. Mixes-off-epochs -20. I used 3x more samples to make 1 batch in case of simultaneous use mixup and cutmix. Also, I tried to use 4x more samples but didn’t get a boost.\n- EMA, EMA-decay 0.9999\n- Fonts as external data with ElasticTransform Augs\n- Drop top 300 hardest samples by probability difference\n- Knowledge distillation. I tried make distillation effnet b5 -&gt; effnet b3. On my CV b3 works exactly the same as b5, but on LB score was much worse\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F407384%2F81a19fd7683aa0ae12ed86236b4d65c9%2FScreenshot_20200317_023419.png?generation=1584401681213225&amp;alt=media\" alt=\"\"></p>\n\n<p>Didn't work:\nany augs, rand augment from imagenet, label smoothing</p>\n\n<p>Didn’t try:\nohem, custom loss, cutmask, gridmask. </p>\n\n<p>Special thanks to <a href=\"/romul0212\">@romul0212</a>  for team participation!</p>\n\n<p>Also thanks HOSTKEY for support with grant program. I used one machine with 4x 1080Ti.</p>\n\n<p>post private LB update:\nLol, didn't expect that much drop :)</p>",
  "messages": [
    {
      "id": "775669",
      "postDate": "03/16/2020 23:40:01",
      "content": "<p>There are no insights into my solution that could ruin the LB.  I think I can share now.</p>\n\n<p>Overview of my solution:\n- 5folds, random split (seems the main problem)\n- Original size - 236,137, Also tried cropping, resizing\n- Apex, O0 opt level (pure fp32). For some reason O1 opt level worsened score\n- Best models: efficientnet b5 noisy student, wsl_resnext101. It’s my first time then efficientnet works! \n- AdamW, initLR 3e-4, minLR 1e-6\n- Warmup 3 epochs + Cosine Decay 200 epochs + Cooldown 10 epochs. Didn’t play much with scheduling.\n- Mixup + Cutmix, Cutmix-prob 1.0, Cutmix-alpha 1.0, Mixup-prob 0.4, Mixup-alpha 0.2. Mixes-off-epochs -20. I used 3x more samples to make 1 batch in case of simultaneous use mixup and cutmix. Also, I tried to use 4x more samples but didn’t get a boost.\n- EMA, EMA-decay 0.9999\n- Fonts as external data with ElasticTransform Augs\n- Drop top 300 hardest samples by probability difference\n- Knowledge distillation. I tried make distillation effnet b5 -&gt; effnet b3. On my CV b3 works exactly the same as b5, but on LB score was much worse\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F407384%2F81a19fd7683aa0ae12ed86236b4d65c9%2FScreenshot_20200317_023419.png?generation=1584401681213225&amp;alt=media\" alt=\"\"></p>\n\n<p>Didn't work:\nany augs, rand augment from imagenet, label smoothing</p>\n\n<p>Didn’t try:\nohem, custom loss, cutmask, gridmask. </p>\n\n<p>Special thanks to <a href=\"/romul0212\">@romul0212</a>  for team participation!</p>\n\n<p>Also thanks HOSTKEY for support with grant program. I used one machine with 4x 1080Ti.</p>\n\n<p>post private LB update:\nLol, didn't expect that much drop :)</p>",
      "rawMarkdown": "There are no insights into my solution that could ruin the LB.  I think I can share now.\n\nOverview of my solution:\n- 5folds, random split (seems the main problem)\n- Original size - 236,137, Also tried cropping, resizing\n- Apex, O0 opt level (pure fp32). For some reason O1 opt level worsened score\n- Best models: efficientnet b5 noisy student, wsl_resnext101. It’s my first time then efficientnet works! \n- AdamW, initLR 3e-4, minLR 1e-6\n- Warmup 3 epochs + Cosine Decay 200 epochs + Cooldown 10 epochs. Didn’t play much with scheduling.\n- Mixup + Cutmix, Cutmix-prob 1.0, Cutmix-alpha 1.0, Mixup-prob 0.4, Mixup-alpha 0.2. Mixes-off-epochs -20. I used 3x more samples to make 1 batch in case of simultaneous use mixup and cutmix. Also, I tried to use 4x more samples but didn’t get a boost.\n- EMA, EMA-decay 0.9999\n- Fonts as external data with ElasticTransform Augs\n- Drop top 300 hardest samples by probability difference\n- Knowledge distillation. I tried make distillation effnet b5 -&gt; effnet b3. On my CV b3 works exactly the same as b5, but on LB score was much worse\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F407384%2F81a19fd7683aa0ae12ed86236b4d65c9%2FScreenshot_20200317_023419.png?generation=1584401681213225&amp;alt=media)\n\nDidn't work:\nany augs, rand augment from imagenet, label smoothing\n\nDidn’t try:\nohem, custom loss, cutmask, gridmask. \n\nSpecial thanks to @romul0212  for team participation!\n\nAlso thanks HOSTKEY for support with grant program. I used one machine with 4x 1080Ti.\n\npost private LB update:\nLol, didn't expect that much drop :)",
      "votes": null
    },
    {
      "id": "775675",
      "postDate": "03/16/2020 23:53:37",
      "content": "<p>Thank you and congrats and good luck private.</p>\n\n<ol>\n<li>What is EMA?</li>\n<li>Can you post where to download the architecture of wsl_resnext101. (Wide resnext?)</li>\n<li>How to cutmix across batches. Why do you even think of this idea? Does it help? If it helps, why?</li>\n<li>Why drop top 300 hardest samples by probability difference? We used OHEM which emphasized the hardest samples. So not sure why dropping hard samples is good? Maybe because they were incorrectly labeled? Did you drop per grapheme_root, or just the top 300 overall? Did you see an improvement?</li>\n</ol>\n\n<p>Congrats, gl</p>",
      "rawMarkdown": "Thank you and congrats and good luck private.\n\n1. What is EMA?\n2. Can you post where to download the architecture of wsl_resnext101. (Wide resnext?)\n3. How to cutmix across batches. Why do you even think of this idea? Does it help? If it helps, why?\n4. Why drop top 300 hardest samples by probability difference? We used OHEM which emphasized the hardest samples. So not sure why dropping hard samples is good? Maybe because they were incorrectly labeled? Did you drop per grapheme_root, or just the top 300 overall? Did you see an improvement?\n\nCongrats, gl",
      "votes": null
    },
    {
      "id": "775679",
      "postDate": "03/16/2020 23:55:53",
      "content": "<p>Right now <strong>4 minutes to go</strong> ... why now, why not earlier! 😄 , Anyway thanks and good luck for the private LB. :) </p>",
      "rawMarkdown": "Right now **4 minutes to go** ... why now, why not earlier! 😄 , Anyway thanks and good luck for the private LB. :)",
      "votes": null
    },
    {
      "id": "775682",
      "postDate": "03/16/2020 23:58:35",
      "content": "<p><a href=\"/drn01z3\">@drn01z3</a> So you pseudo-labeled these images? Very cool!</p>",
      "rawMarkdown": "drn01z3 So you pseudo-labeled these images? Very cool!",
      "votes": null
    },
    {
      "id": "775685",
      "postDate": "03/17/2020 00:02:04",
      "content": "<ol>\n<li>Exponetial Moving Average. It's about avereging model weight during training across all epochs.</li>\n<li><a href=\"https://pytorch.org/hub/facebookresearch_WSL-Images_resnext/\">https://pytorch.org/hub/facebookresearch_WSL-Images_resnext/</a></li>\n<li>I build one batch from 3 batches to boost verity of data</li>\n<li>Top 300 - just good number, not a serious study. It's boost local score and a little bit LB.</li>\n</ol>",
      "rawMarkdown": "1. Exponetial Moving Average. It's about avereging model weight during training across all epochs.\n2. https://pytorch.org/hub/facebookresearch_WSL-Images_resnext/\n3. I build one batch from 3 batches to boost verity of data\n4. Top 300 - just good number, not a serious study. It's boost local score and a little bit LB.",
      "votes": null
    },
    {
      "id": "775692",
      "postDate": "03/17/2020 00:07:14",
      "content": "<p>RIP</p>",
      "rawMarkdown": "RIP",
      "votes": null
    },
    {
      "id": "775695",
      "postDate": "03/17/2020 00:08:46",
      "content": "<p><a href=\"/drn01z3\">@drn01z3</a> really impressive as always, first time I see EMA here. when I read this kind of approaches (complete and elaborated) and then I see the results LB, I always think: was it worth it? (I think yes haha)</p>",
      "rawMarkdown": "drn01z3 really impressive as always, first time I see EMA here. when I read this kind of approaches (complete and elaborated) and then I see the results LB, I always think: was it worth it? (I think yes haha)",
      "votes": null
    },
    {
      "id": "775700",
      "postDate": "03/17/2020 00:10:06",
      "content": "<p>Well, actually not. Knowledge distillation is not like pseudo-labeled. Also I didn't use external data (except fonts) to make true pseudo-labeled</p>",
      "rawMarkdown": "Well, actually not. Knowledge distillation is not like pseudo-labeled. Also I didn't use external data (except fonts) to make true pseudo-labeled",
      "votes": null
    },
    {
      "id": "775708",
      "postDate": "03/17/2020 00:16:26",
      "content": "<p>Thanks for the kind words! I am deeply convinced that it was worth it. I first tried distillation. I figured out the correct implementation of mixup / cutmix. And also for the first time, the effnet performed very well.</p>",
      "rawMarkdown": "Thanks for the kind words! I am deeply convinced that it was worth it. I first tried distillation. I figured out the correct implementation of mixup / cutmix. And also for the first time, the effnet performed very well.",
      "votes": null
    },
    {
      "id": "775718",
      "postDate": "03/17/2020 00:24:04",
      "content": "<p>Thanks for sharing. Any idea about the huge gap between public and private LB? </p>",
      "rawMarkdown": "Thanks for sharing. Any idea about the huge gap between public and private LB?",
      "votes": null
    },
    {
      "id": "775723",
      "postDate": "03/17/2020 00:27:56",
      "content": "<p>Unseen combination, I guess. Better to ask guys from the LB top :)</p>",
      "rawMarkdown": "Unseen combination, I guess. Better to ask guys from the LB top :)",
      "votes": null
    },
    {
      "id": "775725",
      "postDate": "03/17/2020 00:28:29",
      "content": "<p>\"Why drop top 300 hardest samples by probability difference? We used OHEM which emphasized the hardest samples\"</p>\n\n<p>i am curious about this as well. is there comparsion results with and without dropping samples?</p>\n\n<p>it is possible that dropping samples is better than OHEM if these dropped samples are outliers</p>",
      "rawMarkdown": "\"Why drop top 300 hardest samples by probability difference? We used OHEM which emphasized the hardest samples\"\n\ni am curious about this as well. is there comparsion results with and without dropping samples?\n\nit is possible that dropping samples is better than OHEM if these dropped samples are outliers",
      "votes": null
    },
    {
      "id": "775727",
      "postDate": "03/17/2020 00:30:59",
      "content": "<p>Can you share any links to learn more about knowledge distillation?</p>",
      "rawMarkdown": "Can you share any links to learn more about knowledge distillation?",
      "votes": null
    },
    {
      "id": "775729",
      "postDate": "03/17/2020 00:32:06",
      "content": "<p>Did you train with TPUs for EfficientNet?</p>",
      "rawMarkdown": "Did you train with TPUs for EfficientNet?",
      "votes": null
    },
    {
      "id": "775732",
      "postDate": "03/17/2020 00:33:29",
      "content": "<p>I think this is the orignal paper about knowledge distillation: <a href=\"https://arxiv.org/abs/1503.02531\">https://arxiv.org/abs/1503.02531</a></p>",
      "rawMarkdown": "I think this is the orignal paper about knowledge distillation: https://arxiv.org/abs/1503.02531",
      "votes": null
    },
    {
      "id": "775736",
      "postDate": "03/17/2020 00:35:00",
      "content": "<p>I checked this on same folds using effnet b3 models. Dropping samples makes score better localy and on LB. I think it's almost impossible to collect 100% pure dataset. So dropping the hardest samples is about remove some noise from data.</p>",
      "rawMarkdown": "I checked this on same folds using effnet b3 models. Dropping samples makes score better localy and on LB. I think it's almost impossible to collect 100% pure dataset. So dropping the hardest samples is about remove some noise from data.",
      "votes": null
    },
    {
      "id": "775746",
      "postDate": "03/17/2020 00:36:17",
      "content": "<p>I didn't use TPU. Just 8x Tesla V100 and 7x 2080Ti. And with HOSTKEY GRANT Programm 4x 1080 Ti.</p>",
      "rawMarkdown": "I didn't use TPU. Just 8x Tesla V100 and 7x 2080Ti. And with HOSTKEY GRANT Programm 4x 1080 Ti.",
      "votes": null
    },
    {
      "id": "776180",
      "postDate": "03/17/2020 07:18:33",
      "content": "<p>i begin to re implement your solution now. if possible, could you share your code, log file (training curve), data split. Thanks in advance!</p>",
      "rawMarkdown": "i begin to re implement your solution now. if possible, could you share your code, log file (training curve), data split. Thanks in advance!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 775675,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "03/16/2020 23:53:37",
      "content": "<p>Thank you and congrats and good luck private.</p>\n\n<ol>\n<li>What is EMA?</li>\n<li>Can you post where to download the architecture of wsl_resnext101. (Wide resnext?)</li>\n<li>How to cutmix across batches. Why do you even think of this idea? Does it help? If it helps, why?</li>\n<li>Why drop top 300 hardest samples by probability difference? We used OHEM which emphasized the hardest samples. So not sure why dropping hard samples is good? Maybe because they were incorrectly labeled? Did you drop per grapheme_root, or just the top 300 overall? Did you see an improvement?</li>\n</ol>\n\n<p>Congrats, gl</p>",
      "votes": null,
      "replies": [
        {
          "id": 775685,
          "author_name": "drn01z3",
          "author_url": "",
          "post_date": "03/17/2020 00:02:04",
          "content": "<ol>\n<li>Exponetial Moving Average. It's about avereging model weight during training across all epochs.</li>\n<li><a href=\"https://pytorch.org/hub/facebookresearch_WSL-Images_resnext/\">https://pytorch.org/hub/facebookresearch_WSL-Images_resnext/</a></li>\n<li>I build one batch from 3 batches to boost verity of data</li>\n<li>Top 300 - just good number, not a serious study. It's boost local score and a little bit LB.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775725,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/17/2020 00:28:29",
          "content": "<p>\"Why drop top 300 hardest samples by probability difference? We used OHEM which emphasized the hardest samples\"</p>\n\n<p>i am curious about this as well. is there comparsion results with and without dropping samples?</p>\n\n<p>it is possible that dropping samples is better than OHEM if these dropped samples are outliers</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775736,
          "author_name": "drn01z3",
          "author_url": "",
          "post_date": "03/17/2020 00:35:00",
          "content": "<p>I checked this on same folds using effnet b3 models. Dropping samples makes score better localy and on LB. I think it's almost impossible to collect 100% pure dataset. So dropping the hardest samples is about remove some noise from data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775679,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "03/16/2020 23:55:53",
      "content": "<p>Right now <strong>4 minutes to go</strong> ... why now, why not earlier! 😄 , Anyway thanks and good luck for the private LB. :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 775682,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "03/16/2020 23:58:35",
      "content": "<p><a href=\"/drn01z3\">@drn01z3</a> So you pseudo-labeled these images? Very cool!</p>",
      "votes": null,
      "replies": [
        {
          "id": 775700,
          "author_name": "drn01z3",
          "author_url": "",
          "post_date": "03/17/2020 00:10:06",
          "content": "<p>Well, actually not. Knowledge distillation is not like pseudo-labeled. Also I didn't use external data (except fonts) to make true pseudo-labeled</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775692,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "03/17/2020 00:07:14",
      "content": "<p>RIP</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 775695,
      "author_name": "jesucristo",
      "author_url": "",
      "post_date": "03/17/2020 00:08:46",
      "content": "<p><a href=\"/drn01z3\">@drn01z3</a> really impressive as always, first time I see EMA here. when I read this kind of approaches (complete and elaborated) and then I see the results LB, I always think: was it worth it? (I think yes haha)</p>",
      "votes": null,
      "replies": [
        {
          "id": 775708,
          "author_name": "drn01z3",
          "author_url": "",
          "post_date": "03/17/2020 00:16:26",
          "content": "<p>Thanks for the kind words! I am deeply convinced that it was worth it. I first tried distillation. I figured out the correct implementation of mixup / cutmix. And also for the first time, the effnet performed very well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775718,
      "author_name": "gxygomes",
      "author_url": "",
      "post_date": "03/17/2020 00:24:04",
      "content": "<p>Thanks for sharing. Any idea about the huge gap between public and private LB? </p>",
      "votes": null,
      "replies": [
        {
          "id": 775723,
          "author_name": "drn01z3",
          "author_url": "",
          "post_date": "03/17/2020 00:27:56",
          "content": "<p>Unseen combination, I guess. Better to ask guys from the LB top :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775727,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "03/17/2020 00:30:59",
      "content": "<p>Can you share any links to learn more about knowledge distillation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 775732,
          "author_name": "gxygomes",
          "author_url": "",
          "post_date": "03/17/2020 00:33:29",
          "content": "<p>I think this is the orignal paper about knowledge distillation: <a href=\"https://arxiv.org/abs/1503.02531\">https://arxiv.org/abs/1503.02531</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775729,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "03/17/2020 00:32:06",
      "content": "<p>Did you train with TPUs for EfficientNet?</p>",
      "votes": null,
      "replies": [
        {
          "id": 775746,
          "author_name": "drn01z3",
          "author_url": "",
          "post_date": "03/17/2020 00:36:17",
          "content": "<p>I didn't use TPU. Just 8x Tesla V100 and 7x 2080Ti. And with HOSTKEY GRANT Programm 4x 1080 Ti.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776180,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/17/2020 07:18:33",
      "content": "<p>i begin to re implement your solution now. if possible, could you share your code, log file (training curve), data split. Thanks in advance!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "775669": "There are no insights into my solution that could ruin the LB.  I think I can share now.\n\nOverview of my solution:\n- 5folds, random split (seems the main problem)\n- Original size - 236,137, Also tried cropping, resizing\n- Apex, O0 opt level (pure fp32). For some reason O1 opt level worsened score\n- Best models: efficientnet b5 noisy student, wsl_resnext101. It’s my first time then efficientnet works! \n- AdamW, initLR 3e-4, minLR 1e-6\n- Warmup 3 epochs + Cosine Decay 200 epochs + Cooldown 10 epochs. Didn’t play much with scheduling.\n- Mixup + Cutmix, Cutmix-prob 1.0, Cutmix-alpha 1.0, Mixup-prob 0.4, Mixup-alpha 0.2. Mixes-off-epochs -20. I used 3x more samples to make 1 batch in case of simultaneous use mixup and cutmix. Also, I tried to use 4x more samples but didn’t get a boost.\n- EMA, EMA-decay 0.9999\n- Fonts as external data with ElasticTransform Augs\n- Drop top 300 hardest samples by probability difference\n- Knowledge distillation. I tried make distillation effnet b5 -&gt; effnet b3. On my CV b3 works exactly the same as b5, but on LB score was much worse\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F407384%2F81a19fd7683aa0ae12ed86236b4d65c9%2FScreenshot_20200317_023419.png?generation=1584401681213225&amp;alt=media)\n\nDidn't work:\nany augs, rand augment from imagenet, label smoothing\n\nDidn’t try:\nohem, custom loss, cutmask, gridmask. \n\nSpecial thanks to @romul0212  for team participation!\n\nAlso thanks HOSTKEY for support with grant program. I used one machine with 4x 1080Ti.\n\npost private LB update:\nLol, didn't expect that much drop :)",
    "775675": "Thank you and congrats and good luck private.\n\n1. What is EMA?\n2. Can you post where to download the architecture of wsl_resnext101. (Wide resnext?)\n3. How to cutmix across batches. Why do you even think of this idea? Does it help? If it helps, why?\n4. Why drop top 300 hardest samples by probability difference? We used OHEM which emphasized the hardest samples. So not sure why dropping hard samples is good? Maybe because they were incorrectly labeled? Did you drop per grapheme_root, or just the top 300 overall? Did you see an improvement?\n\nCongrats, gl",
    "775679": "Right now **4 minutes to go** ... why now, why not earlier! 😄 , Anyway thanks and good luck for the private LB. :)",
    "775682": "drn01z3 So you pseudo-labeled these images? Very cool!",
    "775685": "1. Exponetial Moving Average. It's about avereging model weight during training across all epochs.\n2. https://pytorch.org/hub/facebookresearch_WSL-Images_resnext/\n3. I build one batch from 3 batches to boost verity of data\n4. Top 300 - just good number, not a serious study. It's boost local score and a little bit LB.",
    "775692": "RIP",
    "775695": "drn01z3 really impressive as always, first time I see EMA here. when I read this kind of approaches (complete and elaborated) and then I see the results LB, I always think: was it worth it? (I think yes haha)",
    "775700": "Well, actually not. Knowledge distillation is not like pseudo-labeled. Also I didn't use external data (except fonts) to make true pseudo-labeled",
    "775708": "Thanks for the kind words! I am deeply convinced that it was worth it. I first tried distillation. I figured out the correct implementation of mixup / cutmix. And also for the first time, the effnet performed very well.",
    "775718": "Thanks for sharing. Any idea about the huge gap between public and private LB?",
    "775723": "Unseen combination, I guess. Better to ask guys from the LB top :)",
    "775725": "\"Why drop top 300 hardest samples by probability difference? We used OHEM which emphasized the hardest samples\"\n\ni am curious about this as well. is there comparsion results with and without dropping samples?\n\nit is possible that dropping samples is better than OHEM if these dropped samples are outliers",
    "775727": "Can you share any links to learn more about knowledge distillation?",
    "775729": "Did you train with TPUs for EfficientNet?",
    "775732": "I think this is the orignal paper about knowledge distillation: https://arxiv.org/abs/1503.02531",
    "775736": "I checked this on same folds using effnet b3 models. Dropping samples makes score better localy and on LB. I think it's almost impossible to collect 100% pure dataset. So dropping the hardest samples is about remove some noise from data.",
    "775746": "I didn't use TPU. Just 8x Tesla V100 and 7x 2080Ti. And with HOSTKEY GRANT Programm 4x 1080 Ti.",
    "776180": "i begin to re implement your solution now. if possible, could you share your code, log file (training curve), data split. Thanks in advance!"
  },
  "source": "meta"
}