{
  "id": 136056,
  "title": "21st Solution using only Kaggle Kernels",
  "url": "/competitions/bengaliai-cv19/writeups/21st-solution-using-only-kaggle-kernels",
  "author_name": "",
  "post_date": "2020-03-17T07:51:27.257110700Z",
  "votes": 58,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I would like to thank <a href=\"/seesee\">@seesee</a> a lot, because he was very generous to share his notebooks: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134161\">https://www.kaggle.com/c/bengaliai-cv19/discussion/134161</a> I was not going to join this competition because I have no GPU, but in the last week, I have decided to join after seeing these nice kernels. I could only use 2x30 TPU hours for this competition.</p>\n\n<p>The baseline was scoring 0.9708 on the public LB. Here is what I have done on top of that:\n* I was aware that <strong>random split</strong> for train and validation does not represent the split for the real test set but I was lazy to change it. I just kept it in mind.\n* Every class contributes to the metric equally. Therefore, I have <strong>multiplied the predictions by inverse class frequencies</strong> before getting the maximum. This gave 0.007 boost on Public LB. (around 0.003 boost for the validation set)\n* I have changed the backbone model from <strong>Efficientnet B3 to B4</strong>. This gave 0.002 boost.\n* I have changed the augmentation from mixup to <strong>cutmix ** using the implementation here: <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\">https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu</a> I have only updated it so that it does the cuts as rectangles rather than squares (and fixed a small bug, I believe). Thanks <a href=\"/cdeotte\">@cdeotte</a> for this nice kernel. cutmix did not improve my LB but improved the local validation.\n* I have changed the learning scheme. **Decreasing learning rates gradually</strong> gave me 0.007 boost. I could make use of it more but TPU run-time limit was 3 hours. If I had higher limit, I could end up in gold zone. Or maybe the opposite, having an underfit model gave me a shake-up.\n* I have ran the model <strong>on whole data</strong>, rather than 80% train split. This gave 0.001 boost.\n* I have <strong>blended the last two runs</strong> with different seeds. This gave me 0.001 boost.\n* Considering that my split was not representative and having more benefit from multiplying by inverse class frequencies than expected, I have decided to exploit a bit more by weight = np.power(weight, 1.2). This improved public LB by 0.0006 but seems to improve private even more. Probably because there are more <strong>unseen graphemes</strong> in the private set.</p>\n\n<p>So it was possible to get 21st place using only Kaggle Kernels in a week. Reading discussions and kernels carefully is enough for getting around this position. For top 10, usually more creativity is needed.</p>",
  "messages": [
    {
      "id": "776211",
      "postDate": "03/17/2020 07:51:27",
      "content": "<p>I would like to thank <a href=\"/seesee\">@seesee</a> a lot, because he was very generous to share his notebooks: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134161\">https://www.kaggle.com/c/bengaliai-cv19/discussion/134161</a> I was not going to join this competition because I have no GPU, but in the last week, I have decided to join after seeing these nice kernels. I could only use 2x30 TPU hours for this competition.</p>\n\n<p>The baseline was scoring 0.9708 on the public LB. Here is what I have done on top of that:\n* I was aware that <strong>random split</strong> for train and validation does not represent the split for the real test set but I was lazy to change it. I just kept it in mind.\n* Every class contributes to the metric equally. Therefore, I have <strong>multiplied the predictions by inverse class frequencies</strong> before getting the maximum. This gave 0.007 boost on Public LB. (around 0.003 boost for the validation set)\n* I have changed the backbone model from <strong>Efficientnet B3 to B4</strong>. This gave 0.002 boost.\n* I have changed the augmentation from mixup to <strong>cutmix ** using the implementation here: <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu\">https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu</a> I have only updated it so that it does the cuts as rectangles rather than squares (and fixed a small bug, I believe). Thanks <a href=\"/cdeotte\">@cdeotte</a> for this nice kernel. cutmix did not improve my LB but improved the local validation.\n* I have changed the learning scheme. **Decreasing learning rates gradually</strong> gave me 0.007 boost. I could make use of it more but TPU run-time limit was 3 hours. If I had higher limit, I could end up in gold zone. Or maybe the opposite, having an underfit model gave me a shake-up.\n* I have ran the model <strong>on whole data</strong>, rather than 80% train split. This gave 0.001 boost.\n* I have <strong>blended the last two runs</strong> with different seeds. This gave me 0.001 boost.\n* Considering that my split was not representative and having more benefit from multiplying by inverse class frequencies than expected, I have decided to exploit a bit more by weight = np.power(weight, 1.2). This improved public LB by 0.0006 but seems to improve private even more. Probably because there are more <strong>unseen graphemes</strong> in the private set.</p>\n\n<p>So it was possible to get 21st place using only Kaggle Kernels in a week. Reading discussions and kernels carefully is enough for getting around this position. For top 10, usually more creativity is needed.</p>",
      "rawMarkdown": "I would like to thank @seesee a lot, because he was very generous to share his notebooks: https://www.kaggle.com/c/bengaliai-cv19/discussion/134161 I was not going to join this competition because I have no GPU, but in the last week, I have decided to join after seeing these nice kernels. I could only use 2x30 TPU hours for this competition.\n\nThe baseline was scoring 0.9708 on the public LB. Here is what I have done on top of that:\n* I was aware that **random split** for train and validation does not represent the split for the real test set but I was lazy to change it. I just kept it in mind.\n* Every class contributes to the metric equally. Therefore, I have **multiplied the predictions by inverse class frequencies** before getting the maximum. This gave 0.007 boost on Public LB. (around 0.003 boost for the validation set)\n* I have changed the backbone model from **Efficientnet B3 to B4**. This gave 0.002 boost.\n* I have changed the augmentation from mixup to **cutmix ** using the implementation here: https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu I have only updated it so that it does the cuts as rectangles rather than squares (and fixed a small bug, I believe). Thanks @cdeotte for this nice kernel. cutmix did not improve my LB but improved the local validation.\n* I have changed the learning scheme. **Decreasing learning rates gradually** gave me 0.007 boost. I could make use of it more but TPU run-time limit was 3 hours. If I had higher limit, I could end up in gold zone. Or maybe the opposite, having an underfit model gave me a shake-up.\n* I have ran the model **on whole data**, rather than 80% train split. This gave 0.001 boost.\n* I have **blended the last two runs** with different seeds. This gave me 0.001 boost.\n* Considering that my split was not representative and having more benefit from multiplying by inverse class frequencies than expected, I have decided to exploit a bit more by weight = np.power(weight, 1.2). This improved public LB by 0.0006 but seems to improve private even more. Probably because there are more **unseen graphemes** in the private set.\n\nSo it was possible to get 21st place using only Kaggle Kernels in a week. Reading discussions and kernels carefully is enough for getting around this position. For top 10, usually more creativity is needed.",
      "votes": null
    },
    {
      "id": "776229",
      "postDate": "03/17/2020 08:08:05",
      "content": "<blockquote>\n  <p>it was possible to get 21st place using only Kaggle Kernels in a week.</p>\n</blockquote>\n\n<p>Very very impressive. The fact that you did it with only kaggle-kernels is strong message to those who think that you need heavy compute power to do well in kaggle(vision) competitions.</p>",
      "rawMarkdown": "&gt; it was possible to get 21st place using only Kaggle Kernels in a week.\n\nVery very impressive. The fact that you did it with only kaggle-kernels is strong message to those who think that you need heavy compute power to do well in kaggle(vision) competitions.",
      "votes": null
    },
    {
      "id": "776230",
      "postDate": "03/17/2020 08:15:39",
      "content": "<p>Thanks. Indeed. But I was lucky that I started from a great baseline by <a href=\"/seesee\">@seesee</a> </p>",
      "rawMarkdown": "Thanks. Indeed. But I was lucky that I started from a great baseline by @seesee",
      "votes": null
    },
    {
      "id": "776240",
      "postDate": "03/17/2020 08:27:14",
      "content": "<p>Congratulations.</p>\n\n<p>Thanks for clarifying 'Kaggle Kernels' it helps in future for people like me who do not have dedicated GPU. it builds confidence for future comps</p>",
      "rawMarkdown": "Congratulations.\n\nThanks for clarifying 'Kaggle Kernels' it helps in future for people like me who do not have dedicated GPU. it builds confidence for future comps",
      "votes": null
    },
    {
      "id": "776247",
      "postDate": "03/17/2020 08:29:54",
      "content": "<p>🙏 Congratulations! Thanks for sharing. I too only used kaggle kernel and gained from the shakeup but could not make it to medals.  😔 </p>",
      "rawMarkdown": "🙏 Congratulations! Thanks for sharing. I too only used kaggle kernel and gained from the shakeup but could not make it to medals.  😔",
      "votes": null
    },
    {
      "id": "776248",
      "postDate": "03/17/2020 08:30:52",
      "content": "<p>Congrats, that's an awesome solution! I am glad that my notebooks were helpful.</p>\n\n<blockquote>\n  <p>If I had higher limit, I could end up in gold zone. Or maybe the opposite, having an underfit model gave me a shake-up.</p>\n</blockquote>\n\n<p>Probably gold ;). I think that the post-processing made the difference.</p>",
      "rawMarkdown": "Congrats, that's an awesome solution! I am glad that my notebooks were helpful.\n\n&gt; If I had higher limit, I could end up in gold zone. Or maybe the opposite, having an underfit model gave me a shake-up.\n\nProbably gold ;). I think that the post-processing made the difference.",
      "votes": null
    },
    {
      "id": "776252",
      "postDate": "03/17/2020 08:36:05",
      "content": "<p>&gt; Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum</p>\n\n<p>Lesson learnt, thanks</p>",
      "rawMarkdown": "&gt; Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum\n\nLesson learnt, thanks",
      "votes": null
    },
    {
      "id": "776282",
      "postDate": "03/17/2020 09:08:12",
      "content": "<p>Awesome.  </p>\n\n<blockquote>\n  <p>Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum.</p>\n</blockquote>\n\n<p>We started looking at this kind of postprocessing the afternoon of last day, and I found a scheme a bit different 2 hours before deadline that worked marvel locally.  Will code and submit to check.</p>",
      "rawMarkdown": "Awesome.  \n\n&gt; Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum.\n\nWe started looking at this kind of postprocessing the afternoon of last day, and I found a scheme a bit different 2 hours before deadline that worked marvel locally.  Will code and submit to check.",
      "votes": null
    },
    {
      "id": "776289",
      "postDate": "03/17/2020 09:13:28",
      "content": "<p>Amazing, also simple, yet effective post processing which is what specifically helped on unseen graphemes.</p>",
      "rawMarkdown": "Amazing, also simple, yet effective post processing which is what specifically helped on unseen graphemes.",
      "votes": null
    },
    {
      "id": "776300",
      "postDate": "03/17/2020 09:31:53",
      "content": "<blockquote>\n  <p>The fact that you did it with only kaggle-kernels is strong message to those who think that you need heavy compute power to do well in kaggle(vision) competitions.</p>\n</blockquote>\n\n<p>True, but having some Turing GPUs don't really hurt. Atleast you can play some good looking games if things turns out bad :) I had trouble setting up TPUs with pytorch. Will try it in some other competition. </p>",
      "rawMarkdown": "&gt; The fact that you did it with only kaggle-kernels is strong message to those who think that you need heavy compute power to do well in kaggle(vision) competitions.\n\nTrue, but having some Turing GPUs don't really hurt. Atleast you can play some good looking games if things turns out bad :) I had trouble setting up TPUs with pytorch. Will try it in some other competition.",
      "votes": null
    },
    {
      "id": "776359",
      "postDate": "03/17/2020 10:30:09",
      "content": "<p>Yep, it's the post-processing that makes the difference. I get +0.026 on private :)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fdf1037144bb4a81cb23323ae5529915c%2Fno-pp.png?generation=1584440814076645&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fb111ab26740105e5606ba1a5157e4e74%2Fahmets-pp.png?generation=1584440884436004&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Yep, it's the post-processing that makes the difference. I get +0.026 on private :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fdf1037144bb4a81cb23323ae5529915c%2Fno-pp.png?generation=1584440814076645&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fb111ab26740105e5606ba1a5157e4e74%2Fahmets-pp.png?generation=1584440884436004&amp;alt=media)",
      "votes": null
    },
    {
      "id": "776368",
      "postDate": "03/17/2020 10:39:45",
      "content": "<p>I was actually assuming that all the top teams were doing the same postprocessing:)</p>",
      "rawMarkdown": "I was actually assuming that all the top teams were doing the same postprocessing:)",
      "votes": null
    },
    {
      "id": "776405",
      "postDate": "03/17/2020 11:15:01",
      "content": "<blockquote>\n  <p>Every class contributes to the metric equally. Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum. This gave 0.007 boost on Public LB. (around 0.003 boost for the validation set)</p>\n</blockquote>\n\n<p>I actually use that rescaling trick a lot at work against unbalanced data.</p>",
      "rawMarkdown": "&gt; Every class contributes to the metric equally. Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum. This gave 0.007 boost on Public LB. (around 0.003 boost for the validation set)\n\nI actually use that rescaling trick a lot at work against unbalanced data.",
      "votes": null
    },
    {
      "id": "776666",
      "postDate": "03/17/2020 14:42:57",
      "content": "<p>Amazing job Ahmet. Congrats. I am so impressed that you did so much in 1 week. Plus you only used Kaggle notebooks. Incredible!</p>",
      "rawMarkdown": "Amazing job Ahmet. Congrats. I am so impressed that you did so much in 1 week. Plus you only used Kaggle notebooks. Incredible!",
      "votes": null
    },
    {
      "id": "776668",
      "postDate": "03/17/2020 14:44:00",
      "content": "<p>I love your post process trick. Our team did a similar thing <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">here</a>. I'm surprised that more teams did not do this. This would have been a much different comp if every team did this post process.</p>",
      "rawMarkdown": "I love your post process trick. Our team did a similar thing [here][1]. I'm surprised that more teams did not do this. This would have been a much different comp if every team did this post process.\n\n[1]: https://www.kaggle.com/c/bengaliai-cv19/discussion/136021",
      "votes": null
    },
    {
      "id": "779654",
      "postDate": "03/19/2020 15:12:27",
      "content": "<p>Thanks for sharing. In the last point, what is the <code>weight</code>?</p>",
      "rawMarkdown": "Thanks for sharing. In the last point, what is the `weight`?",
      "votes": null
    },
    {
      "id": "779666",
      "postDate": "03/19/2020 15:24:08",
      "content": "<p>Inverse class frequencies</p>",
      "rawMarkdown": "Inverse class frequencies",
      "votes": null
    },
    {
      "id": "779688",
      "postDate": "03/19/2020 15:41:33",
      "content": "<p>I see. Thanks. The inverse weight is really smart!</p>",
      "rawMarkdown": "I see. Thanks. The inverse weight is really smart!",
      "votes": null
    },
    {
      "id": "782155",
      "postDate": "03/22/2020 01:33:30",
      "content": "<p>Cool! Only one week by kaggle TPU, you've done a fantastic job!</p>",
      "rawMarkdown": "Cool! Only one week by kaggle TPU, you've done a fantastic job!",
      "votes": null
    },
    {
      "id": "783530",
      "postDate": "03/23/2020 12:54:53",
      "content": "<p>Amazing job Ahmet. Congrats.</p>",
      "rawMarkdown": "Amazing job Ahmet. Congrats.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 776229,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "03/17/2020 08:08:05",
      "content": "<blockquote>\n  <p>it was possible to get 21st place using only Kaggle Kernels in a week.</p>\n</blockquote>\n\n<p>Very very impressive. The fact that you did it with only kaggle-kernels is strong message to those who think that you need heavy compute power to do well in kaggle(vision) competitions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 776230,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "03/17/2020 08:15:39",
          "content": "<p>Thanks. Indeed. But I was lucky that I started from a great baseline by <a href=\"/seesee\">@seesee</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 776300,
          "author_name": "utsavnandi",
          "author_url": "",
          "post_date": "03/17/2020 09:31:53",
          "content": "<blockquote>\n  <p>The fact that you did it with only kaggle-kernels is strong message to those who think that you need heavy compute power to do well in kaggle(vision) competitions.</p>\n</blockquote>\n\n<p>True, but having some Turing GPUs don't really hurt. Atleast you can play some good looking games if things turns out bad :) I had trouble setting up TPUs with pytorch. Will try it in some other competition. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776240,
      "author_name": "mks2192",
      "author_url": "",
      "post_date": "03/17/2020 08:27:14",
      "content": "<p>Congratulations.</p>\n\n<p>Thanks for clarifying 'Kaggle Kernels' it helps in future for people like me who do not have dedicated GPU. it builds confidence for future comps</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 776247,
      "author_name": "chandraroy",
      "author_url": "",
      "post_date": "03/17/2020 08:29:54",
      "content": "<p>🙏 Congratulations! Thanks for sharing. I too only used kaggle kernel and gained from the shakeup but could not make it to medals.  😔 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 776248,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "03/17/2020 08:30:52",
      "content": "<p>Congrats, that's an awesome solution! I am glad that my notebooks were helpful.</p>\n\n<blockquote>\n  <p>If I had higher limit, I could end up in gold zone. Or maybe the opposite, having an underfit model gave me a shake-up.</p>\n</blockquote>\n\n<p>Probably gold ;). I think that the post-processing made the difference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 776252,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "03/17/2020 08:36:05",
      "content": "<p>&gt; Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum</p>\n\n<p>Lesson learnt, thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 776282,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/17/2020 09:08:12",
      "content": "<p>Awesome.  </p>\n\n<blockquote>\n  <p>Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum.</p>\n</blockquote>\n\n<p>We started looking at this kind of postprocessing the afternoon of last day, and I found a scheme a bit different 2 hours before deadline that worked marvel locally.  Will code and submit to check.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 776289,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "03/17/2020 09:13:28",
      "content": "<p>Amazing, also simple, yet effective post processing which is what specifically helped on unseen graphemes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 776359,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "03/17/2020 10:30:09",
      "content": "<p>Yep, it's the post-processing that makes the difference. I get +0.026 on private :)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fdf1037144bb4a81cb23323ae5529915c%2Fno-pp.png?generation=1584440814076645&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fb111ab26740105e5606ba1a5157e4e74%2Fahmets-pp.png?generation=1584440884436004&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 776368,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "03/17/2020 10:39:45",
          "content": "<p>I was actually assuming that all the top teams were doing the same postprocessing:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776405,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "03/17/2020 11:15:01",
      "content": "<blockquote>\n  <p>Every class contributes to the metric equally. Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum. This gave 0.007 boost on Public LB. (around 0.003 boost for the validation set)</p>\n</blockquote>\n\n<p>I actually use that rescaling trick a lot at work against unbalanced data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 776666,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/17/2020 14:42:57",
      "content": "<p>Amazing job Ahmet. Congrats. I am so impressed that you did so much in 1 week. Plus you only used Kaggle notebooks. Incredible!</p>",
      "votes": null,
      "replies": [
        {
          "id": 776668,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/17/2020 14:44:00",
          "content": "<p>I love your post process trick. Our team did a similar thing <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">here</a>. I'm surprised that more teams did not do this. This would have been a much different comp if every team did this post process.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 779654,
      "author_name": "gxygomes",
      "author_url": "",
      "post_date": "03/19/2020 15:12:27",
      "content": "<p>Thanks for sharing. In the last point, what is the <code>weight</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 779666,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "03/19/2020 15:24:08",
          "content": "<p>Inverse class frequencies</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 779688,
          "author_name": "gxygomes",
          "author_url": "",
          "post_date": "03/19/2020 15:41:33",
          "content": "<p>I see. Thanks. The inverse weight is really smart!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 782155,
      "author_name": "yuanlin08",
      "author_url": "",
      "post_date": "03/22/2020 01:33:30",
      "content": "<p>Cool! Only one week by kaggle TPU, you've done a fantastic job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 783530,
      "author_name": "sanikamal",
      "author_url": "",
      "post_date": "03/23/2020 12:54:53",
      "content": "<p>Amazing job Ahmet. Congrats.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "776211": "I would like to thank @seesee a lot, because he was very generous to share his notebooks: https://www.kaggle.com/c/bengaliai-cv19/discussion/134161 I was not going to join this competition because I have no GPU, but in the last week, I have decided to join after seeing these nice kernels. I could only use 2x30 TPU hours for this competition.\n\nThe baseline was scoring 0.9708 on the public LB. Here is what I have done on top of that:\n* I was aware that **random split** for train and validation does not represent the split for the real test set but I was lazy to change it. I just kept it in mind.\n* Every class contributes to the metric equally. Therefore, I have **multiplied the predictions by inverse class frequencies** before getting the maximum. This gave 0.007 boost on Public LB. (around 0.003 boost for the validation set)\n* I have changed the backbone model from **Efficientnet B3 to B4**. This gave 0.002 boost.\n* I have changed the augmentation from mixup to **cutmix ** using the implementation here: https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu I have only updated it so that it does the cuts as rectangles rather than squares (and fixed a small bug, I believe). Thanks @cdeotte for this nice kernel. cutmix did not improve my LB but improved the local validation.\n* I have changed the learning scheme. **Decreasing learning rates gradually** gave me 0.007 boost. I could make use of it more but TPU run-time limit was 3 hours. If I had higher limit, I could end up in gold zone. Or maybe the opposite, having an underfit model gave me a shake-up.\n* I have ran the model **on whole data**, rather than 80% train split. This gave 0.001 boost.\n* I have **blended the last two runs** with different seeds. This gave me 0.001 boost.\n* Considering that my split was not representative and having more benefit from multiplying by inverse class frequencies than expected, I have decided to exploit a bit more by weight = np.power(weight, 1.2). This improved public LB by 0.0006 but seems to improve private even more. Probably because there are more **unseen graphemes** in the private set.\n\nSo it was possible to get 21st place using only Kaggle Kernels in a week. Reading discussions and kernels carefully is enough for getting around this position. For top 10, usually more creativity is needed.",
    "776229": "&gt; it was possible to get 21st place using only Kaggle Kernels in a week.\n\nVery very impressive. The fact that you did it with only kaggle-kernels is strong message to those who think that you need heavy compute power to do well in kaggle(vision) competitions.",
    "776230": "Thanks. Indeed. But I was lucky that I started from a great baseline by @seesee",
    "776240": "Congratulations.\n\nThanks for clarifying 'Kaggle Kernels' it helps in future for people like me who do not have dedicated GPU. it builds confidence for future comps",
    "776247": "🙏 Congratulations! Thanks for sharing. I too only used kaggle kernel and gained from the shakeup but could not make it to medals.  😔",
    "776248": "Congrats, that's an awesome solution! I am glad that my notebooks were helpful.\n\n&gt; If I had higher limit, I could end up in gold zone. Or maybe the opposite, having an underfit model gave me a shake-up.\n\nProbably gold ;). I think that the post-processing made the difference.",
    "776252": "&gt; Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum\n\nLesson learnt, thanks",
    "776282": "Awesome.  \n\n&gt; Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum.\n\nWe started looking at this kind of postprocessing the afternoon of last day, and I found a scheme a bit different 2 hours before deadline that worked marvel locally.  Will code and submit to check.",
    "776289": "Amazing, also simple, yet effective post processing which is what specifically helped on unseen graphemes.",
    "776300": "&gt; The fact that you did it with only kaggle-kernels is strong message to those who think that you need heavy compute power to do well in kaggle(vision) competitions.\n\nTrue, but having some Turing GPUs don't really hurt. Atleast you can play some good looking games if things turns out bad :) I had trouble setting up TPUs with pytorch. Will try it in some other competition.",
    "776359": "Yep, it's the post-processing that makes the difference. I get +0.026 on private :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fdf1037144bb4a81cb23323ae5529915c%2Fno-pp.png?generation=1584440814076645&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fb111ab26740105e5606ba1a5157e4e74%2Fahmets-pp.png?generation=1584440884436004&amp;alt=media)",
    "776368": "I was actually assuming that all the top teams were doing the same postprocessing:)",
    "776405": "&gt; Every class contributes to the metric equally. Therefore, I have multiplied the predictions by inverse class frequencies before getting the maximum. This gave 0.007 boost on Public LB. (around 0.003 boost for the validation set)\n\nI actually use that rescaling trick a lot at work against unbalanced data.",
    "776666": "Amazing job Ahmet. Congrats. I am so impressed that you did so much in 1 week. Plus you only used Kaggle notebooks. Incredible!",
    "776668": "I love your post process trick. Our team did a similar thing [here][1]. I'm surprised that more teams did not do this. This would have been a much different comp if every team did this post process.\n\n[1]: https://www.kaggle.com/c/bengaliai-cv19/discussion/136021",
    "779654": "Thanks for sharing. In the last point, what is the `weight`?",
    "779666": "Inverse class frequencies",
    "779688": "I see. Thanks. The inverse weight is really smart!",
    "782155": "Cool! Only one week by kaggle TPU, you've done a fantastic job!",
    "783530": "Amazing job Ahmet. Congrats."
  },
  "source": "meta"
}