{
  "id": 123025,
  "title": "fast.ai starter kernel now available [0.964 LB]",
  "url": "/competitions/bengaliai-cv19/discussion/123025",
  "author_name": "",
  "post_date": "2019-12-24T07:16:18.078190800Z",
  "votes": 27,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Greetings, today I have prepared 3 kernels to get started with this competition, which are publicly available now. First, it is <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">image preprocessing kernel</a>, where images with graphemes  are cropped (with some adjustment) and saved as 128x128 images resized with keeping the grapheme aspect ratio. Image stats for this preprocessed images are also available. Next, it is <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-0-960-lb\"> Grapheme fast.ai starter kernel</a>, where the details of the model and training pipeline are provided. Finally, it is the <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference\">inference kernel</a> that got 0.960 LB. I hope these kernels will provide you a clean data and a strong starting point to begin with this competition. Good luck and have fun.</p>\n\n<p>Update: *<em>5-fold can give ~0.966 LB *</em>(while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).</p>\n\n<p>Update: <strong>MixUp gives ~0.003 consistent boost</strong>.  I have modified fast.ai code to make MixUp implementation compatible with the competition data, and the new version of my kernel got 0.9639 LB score for a single fold. 5-fold setup gives  ~0.968, but I do not want to post code (and weights) that cannot be trained within a single kernel run at kaggle. (I decided to disclose this since <a href=\"/hengck23\">@hengck23</a> was planning to post an <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">updated version of his code with MixUp</a>. So I just wanted to make sure that people who are using fast.ai also could apply MixUp). \nHappy New Year! All the best in 2020.</p>",
  "messages": [
    {
      "id": "702030",
      "postDate": "12/24/2019 07:16:18",
      "content": "<p>Greetings, today I have prepared 3 kernels to get started with this competition, which are publicly available now. First, it is <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">image preprocessing kernel</a>, where images with graphemes  are cropped (with some adjustment) and saved as 128x128 images resized with keeping the grapheme aspect ratio. Image stats for this preprocessed images are also available. Next, it is <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-0-960-lb\"> Grapheme fast.ai starter kernel</a>, where the details of the model and training pipeline are provided. Finally, it is the <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference\">inference kernel</a> that got 0.960 LB. I hope these kernels will provide you a clean data and a strong starting point to begin with this competition. Good luck and have fun.</p>\n\n<p>Update: *<em>5-fold can give ~0.966 LB *</em>(while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).</p>\n\n<p>Update: <strong>MixUp gives ~0.003 consistent boost</strong>.  I have modified fast.ai code to make MixUp implementation compatible with the competition data, and the new version of my kernel got 0.9639 LB score for a single fold. 5-fold setup gives  ~0.968, but I do not want to post code (and weights) that cannot be trained within a single kernel run at kaggle. (I decided to disclose this since <a href=\"/hengck23\">@hengck23</a> was planning to post an <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">updated version of his code with MixUp</a>. So I just wanted to make sure that people who are using fast.ai also could apply MixUp). \nHappy New Year! All the best in 2020.</p>",
      "rawMarkdown": "Greetings, today I have prepared 3 kernels to get started with this competition, which are publicly available now. First, it is [image preprocessing kernel](https://www.kaggle.com/iafoss/image-preprocessing-128x128), where images with graphemes  are cropped (with some adjustment) and saved as 128x128 images resized with keeping the grapheme aspect ratio. Image stats for this preprocessed images are also available. Next, it is [ Grapheme fast.ai starter kernel](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-0-960-lb), where the details of the model and training pipeline are provided. Finally, it is the [inference kernel](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference) that got 0.960 LB. I hope these kernels will provide you a clean data and a strong starting point to begin with this competition. Good luck and have fun.\n\nUpdate: **5-fold can give ~0.966 LB **(while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).\n\nUpdate: **MixUp gives ~0.003 consistent boost**.  I have modified fast.ai code to make MixUp implementation compatible with the competition data, and the new version of my kernel got 0.9639 LB score for a single fold. 5-fold setup gives  ~0.968, but I do not want to post code (and weights) that cannot be trained within a single kernel run at kaggle. (I decided to disclose this since @hengck23 was planning to post an [updated version of his code with MixUp](https://www.kaggle.com/c/bengaliai-cv19/discussion/123757). So I just wanted to make sure that people who are using fast.ai also could apply MixUp). \nHappy New Year! All the best in 2020.",
      "votes": null
    },
    {
      "id": "702140",
      "postDate": "12/24/2019 10:09:34",
      "content": "<p>There is private data in it so can not be used as it is. though there is one easy way to change the activation function to something else than Mish then the kernel can be run. You have done great work. Thanks for the kernel. It will surely help in getting up to speed. </p>",
      "rawMarkdown": "There is private data in it so can not be used as it is. though there is one easy way to change the activation function to something else than Mish then the kernel can be run. You have done great work. Thanks for the kernel. It will surely help in getting up to speed.",
      "votes": null
    },
    {
      "id": "702343",
      "postDate": "12/24/2019 15:24:48",
      "content": "<p>I didn't realize that utility scripts are not shared, so it is fixed now, thx for pointing out.</p>",
      "rawMarkdown": "I didn't realize that utility scripts are not shared, so it is fixed now, thx for pointing out.",
      "votes": null
    },
    {
      "id": "705733",
      "postDate": "12/29/2019 11:20:43",
      "content": "<p>Hi,</p>\n\n<p>Thanks for providing the kernels.</p>\n\n<p>I am trying to use fast.ai starter kernel. I created the data following your steps using:\n<code>data = (ImageList.from_df(new, path='/kaggle/input', folder='bengaliai-testing', suffix='.png', cols='image_id', convert_mode='L') \n   .random_split_by_pct(0.2)\n   .label_from_df(cols=['grapheme_root','vowel_diacritic','consonant_diacritic'])\n   .transform(get_transforms(do_flip=False,max_warp=0.1), size=128, padding_mode='zeros')\n   .databunch(bs = 4)).normalize(imagenet_stats)</code></p>\n\n<p>But when I do <code>data.show_batch()</code> I don't see all the three labels for images.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F349898%2F2c6da0244fd8cc82b1d50c4618ad8ea2%2FScreenshot%20(509\" alt=\"\">.png?generation=1577617968533600&amp;alt=media)</p>\n\n<p>My dataframe from which I am trying to create data is as follows:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F349898%2F32706e5e9c9cb499a256bbf1b5d0eab8%2FScreenshot%20(508\" alt=\"\">.png?generation=1577618155515347&amp;alt=media)</p>\n\n<p>Ideally, there should be three labels corresponding to each image, right?\nDo you have any idea why it's happening? Could you please help?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi,\n\nThanks for providing the kernels.\n\nI am trying to use fast.ai starter kernel. I created the data following your steps using:\n`data = (ImageList.from_df(new, path='/kaggle/input', folder='bengaliai-testing', suffix='.png', cols='image_id', convert_mode='L') \n   .random_split_by_pct(0.2)\n   .label_from_df(cols=['grapheme_root','vowel_diacritic','consonant_diacritic'])\n   .transform(get_transforms(do_flip=False,max_warp=0.1), size=128, padding_mode='zeros')\n   .databunch(bs = 4)).normalize(imagenet_stats)`\n\nBut when I do `data.show_batch()` I don't see all the three labels for images.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F349898%2F2c6da0244fd8cc82b1d50c4618ad8ea2%2FScreenshot%20(509).png?generation=1577617968533600&amp;alt=media)\n\nMy dataframe from which I am trying to create data is as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F349898%2F32706e5e9c9cb499a256bbf1b5d0eab8%2FScreenshot%20(508).png?generation=1577618155515347&amp;alt=media)\n\nIdeally, there should be three labels corresponding to each image, right?\nDo you have any idea why it's happening? Could you please help?\n\nThanks",
      "votes": null
    },
    {
      "id": "705875",
      "postDate": "12/29/2019 15:51:56",
      "content": "<p>I think there is a bug in fast.ai  when you try to show data with multiple labels, so just ignore it, or try to fix it by your own if it is really needed. You can check the produced output by <code>x,y = next(iter(data.dls[0]))</code>.</p>",
      "rawMarkdown": "I think there is a bug in fast.ai  when you try to show data with multiple labels, so just ignore it, or try to fix it by your own if it is really needed. You can check the produced output by `x,y = next(iter(data.dls[0]))`.",
      "votes": null
    },
    {
      "id": "705890",
      "postDate": "12/29/2019 16:27:48",
      "content": "<p>Okay. Thanks!</p>",
      "rawMarkdown": "Okay. Thanks!",
      "votes": null
    },
    {
      "id": "706998",
      "postDate": "12/31/2019 06:21:03",
      "content": "<p>Just an update on the method presented here: 5-fold can give ~0.966 LB (while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).</p>",
      "rawMarkdown": "Just an update on the method presented here: 5-fold can give ~0.966 LB (while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).",
      "votes": null
    },
    {
      "id": "707596",
      "postDate": "01/01/2020 05:59:56",
      "content": "<p>MixUp is added to the code.</p>",
      "rawMarkdown": "MixUp is added to the code.",
      "votes": null
    },
    {
      "id": "707611",
      "postDate": "01/01/2020 06:44:15",
      "content": "<p>\"Just an update on the method presented here: 5-fold can give ~0.966 LB (while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).\"</p>\n\n<p>in my experiments, ResNeXt50SE  is slightly better than Densenet121 (in the range of 0.002 in local CV). nevertheless, both results are different, e.g. ResNeXt50SE  is stronger in vowels while Densenet121  is better in root.  Ensemble of may be better</p>",
      "rawMarkdown": "\"Just an update on the method presented here: 5-fold can give ~0.966 LB (while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).\"\n\nin my experiments, ResNeXt50SE  is slightly better than Densenet121 (in the range of 0.002 in local CV). nevertheless, both results are different, e.g. ResNeXt50SE  is stronger in vowels while Densenet121  is better in root.  Ensemble of may be better",
      "votes": null
    },
    {
      "id": "711643",
      "postDate": "01/06/2020 10:46:56",
      "content": "<p>did you calculate the mean/std using cropped or uncropped images?</p>",
      "rawMarkdown": "did you calculate the mean/std using cropped or uncropped images?",
      "votes": null
    },
    {
      "id": "711841",
      "postDate": "01/06/2020 15:17:56",
      "content": "<p>Check <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a> for stats. They are computed on preprocessed images.</p>",
      "rawMarkdown": "Check https://www.kaggle.com/iafoss/image-preprocessing-128x128 for stats. They are computed on preprocessed images.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 702140,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "12/24/2019 10:09:34",
      "content": "<p>There is private data in it so can not be used as it is. though there is one easy way to change the activation function to something else than Mish then the kernel can be run. You have done great work. Thanks for the kernel. It will surely help in getting up to speed. </p>",
      "votes": null,
      "replies": [
        {
          "id": 702343,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/24/2019 15:24:48",
          "content": "<p>I didn't realize that utility scripts are not shared, so it is fixed now, thx for pointing out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 705733,
      "author_name": "romy25",
      "author_url": "",
      "post_date": "12/29/2019 11:20:43",
      "content": "<p>Hi,</p>\n\n<p>Thanks for providing the kernels.</p>\n\n<p>I am trying to use fast.ai starter kernel. I created the data following your steps using:\n<code>data = (ImageList.from_df(new, path='/kaggle/input', folder='bengaliai-testing', suffix='.png', cols='image_id', convert_mode='L') \n   .random_split_by_pct(0.2)\n   .label_from_df(cols=['grapheme_root','vowel_diacritic','consonant_diacritic'])\n   .transform(get_transforms(do_flip=False,max_warp=0.1), size=128, padding_mode='zeros')\n   .databunch(bs = 4)).normalize(imagenet_stats)</code></p>\n\n<p>But when I do <code>data.show_batch()</code> I don't see all the three labels for images.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F349898%2F2c6da0244fd8cc82b1d50c4618ad8ea2%2FScreenshot%20(509\" alt=\"\">.png?generation=1577617968533600&amp;alt=media)</p>\n\n<p>My dataframe from which I am trying to create data is as follows:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F349898%2F32706e5e9c9cb499a256bbf1b5d0eab8%2FScreenshot%20(508\" alt=\"\">.png?generation=1577618155515347&amp;alt=media)</p>\n\n<p>Ideally, there should be three labels corresponding to each image, right?\nDo you have any idea why it's happening? Could you please help?</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 705875,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/29/2019 15:51:56",
          "content": "<p>I think there is a bug in fast.ai  when you try to show data with multiple labels, so just ignore it, or try to fix it by your own if it is really needed. You can check the produced output by <code>x,y = next(iter(data.dls[0]))</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 705890,
          "author_name": "romy25",
          "author_url": "",
          "post_date": "12/29/2019 16:27:48",
          "content": "<p>Okay. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 706998,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "12/31/2019 06:21:03",
      "content": "<p>Just an update on the method presented here: 5-fold can give ~0.966 LB (while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).</p>",
      "votes": null,
      "replies": [
        {
          "id": 707611,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/01/2020 06:44:15",
          "content": "<p>\"Just an update on the method presented here: 5-fold can give ~0.966 LB (while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).\"</p>\n\n<p>in my experiments, ResNeXt50SE  is slightly better than Densenet121 (in the range of 0.002 in local CV). nevertheless, both results are different, e.g. ResNeXt50SE  is stronger in vowels while Densenet121  is better in root.  Ensemble of may be better</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 707596,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "01/01/2020 05:59:56",
      "content": "<p>MixUp is added to the code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 711643,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "01/06/2020 10:46:56",
      "content": "<p>did you calculate the mean/std using cropped or uncropped images?</p>",
      "votes": null,
      "replies": [
        {
          "id": 711841,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "01/06/2020 15:17:56",
          "content": "<p>Check <a href=\"https://www.kaggle.com/iafoss/image-preprocessing-128x128\">https://www.kaggle.com/iafoss/image-preprocessing-128x128</a> for stats. They are computed on preprocessed images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "702030": "Greetings, today I have prepared 3 kernels to get started with this competition, which are publicly available now. First, it is [image preprocessing kernel](https://www.kaggle.com/iafoss/image-preprocessing-128x128), where images with graphemes  are cropped (with some adjustment) and saved as 128x128 images resized with keeping the grapheme aspect ratio. Image stats for this preprocessed images are also available. Next, it is [ Grapheme fast.ai starter kernel](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-0-960-lb), where the details of the model and training pipeline are provided. Finally, it is the [inference kernel](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference) that got 0.960 LB. I hope these kernels will provide you a clean data and a strong starting point to begin with this competition. Good luck and have fun.\n\nUpdate: **5-fold can give ~0.966 LB **(while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).\n\nUpdate: **MixUp gives ~0.003 consistent boost**.  I have modified fast.ai code to make MixUp implementation compatible with the competition data, and the new version of my kernel got 0.9639 LB score for a single fold. 5-fold setup gives  ~0.968, but I do not want to post code (and weights) that cannot be trained within a single kernel run at kaggle. (I decided to disclose this since @hengck23 was planning to post an [updated version of his code with MixUp](https://www.kaggle.com/c/bengaliai-cv19/discussion/123757). So I just wanted to make sure that people who are using fast.ai also could apply MixUp). \nHappy New Year! All the best in 2020.",
    "702140": "There is private data in it so can not be used as it is. though there is one easy way to change the activation function to something else than Mish then the kernel can be run. You have done great work. Thanks for the kernel. It will surely help in getting up to speed.",
    "702343": "I didn't realize that utility scripts are not shared, so it is fixed now, thx for pointing out.",
    "705733": "Hi,\n\nThanks for providing the kernels.\n\nI am trying to use fast.ai starter kernel. I created the data following your steps using:\n`data = (ImageList.from_df(new, path='/kaggle/input', folder='bengaliai-testing', suffix='.png', cols='image_id', convert_mode='L') \n   .random_split_by_pct(0.2)\n   .label_from_df(cols=['grapheme_root','vowel_diacritic','consonant_diacritic'])\n   .transform(get_transforms(do_flip=False,max_warp=0.1), size=128, padding_mode='zeros')\n   .databunch(bs = 4)).normalize(imagenet_stats)`\n\nBut when I do `data.show_batch()` I don't see all the three labels for images.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F349898%2F2c6da0244fd8cc82b1d50c4618ad8ea2%2FScreenshot%20(509).png?generation=1577617968533600&amp;alt=media)\n\nMy dataframe from which I am trying to create data is as follows:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F349898%2F32706e5e9c9cb499a256bbf1b5d0eab8%2FScreenshot%20(508).png?generation=1577618155515347&amp;alt=media)\n\nIdeally, there should be three labels corresponding to each image, right?\nDo you have any idea why it's happening? Could you please help?\n\nThanks",
    "705875": "I think there is a bug in fast.ai  when you try to show data with multiple labels, so just ignore it, or try to fix it by your own if it is really needed. You can check the produced output by `x,y = next(iter(data.dls[0]))`.",
    "705890": "Okay. Thanks!",
    "706998": "Just an update on the method presented here: 5-fold can give ~0.966 LB (while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).",
    "707596": "MixUp is added to the code.",
    "707611": "\"Just an update on the method presented here: 5-fold can give ~0.966 LB (while I used ResNeXt50SE backbone, it performs nearly the same as Densenet121 for a single model).\"\n\nin my experiments, ResNeXt50SE  is slightly better than Densenet121 (in the range of 0.002 in local CV). nevertheless, both results are different, e.g. ResNeXt50SE  is stronger in vowels while Densenet121  is better in root.  Ensemble of may be better",
    "711643": "did you calculate the mean/std using cropped or uncropped images?",
    "711841": "Check https://www.kaggle.com/iafoss/image-preprocessing-128x128 for stats. They are computed on preprocessed images."
  },
  "source": "meta"
}