{
  "id": 128059,
  "title": "Generating More data",
  "url": "/competitions/bengaliai-cv19/discussion/128059",
  "author_name": "",
  "post_date": "2020-01-28T16:22:41.874819200Z",
  "votes": 38,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi everyone. \nRecently I came across this paper <a href=\"https://arxiv.org/pdf/1911.04252v2.pdf\">https://arxiv.org/pdf/1911.04252v2.pdf</a> (please read it for details). Currently its hold record for top score on Imagenet classification </p>\n\n<p>TL DR from this paper::\n```\n1. Train a small network on your training data\n2) Use completely unrelated Images to make predictions and generate Pseudo Labels\n3) Train a bigger student network with your data, and Pseudo Labels </p>\n\n<p><code>``\nUsing this method authors from  paper were able to improve Imagenet score from</code> 86.4 <code>to</code>88.4`</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F8c006f7e5e33cebbd7a72c135fa0b6f2%2FScreen%20Shot%202020-01-28%20at%2011.27.49%20AM.png?generation=1580228955932575&amp;alt=media\" alt=\"\"></p>\n\n<p>I made a simple kernel that give you example how we can generate random Bengali data which can be used for generating pseudo labels. Obviously this kernel can be improved we can also add more publicly available data. </p>\n\n<p>link:  <a href=\"https://www.kaggle.com/drhabib/generating-more-training-data\">https://www.kaggle.com/drhabib/generating-more-training-data</a></p>\n\n<p>Example Images:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F2f5f1760b623d078ef30ad0883b141fb%2FScreen%20Shot%202020-01-28%20at%2011.20.38%20AM.png?generation=1580228534527305&amp;alt=media\" alt=\"\"></p>\n\n<p>P.S my apologize if someone already posted this article and I failed to mention name =) </p>\n\n<p>Val An thanks for the discussion <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808\">https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808</a></p>",
  "messages": [
    {
      "id": "731429",
      "postDate": "01/28/2020 16:22:41",
      "content": "<p>Hi everyone. \nRecently I came across this paper <a href=\"https://arxiv.org/pdf/1911.04252v2.pdf\">https://arxiv.org/pdf/1911.04252v2.pdf</a> (please read it for details). Currently its hold record for top score on Imagenet classification </p>\n\n<p>TL DR from this paper::\n```\n1. Train a small network on your training data\n2) Use completely unrelated Images to make predictions and generate Pseudo Labels\n3) Train a bigger student network with your data, and Pseudo Labels </p>\n\n<p><code>``\nUsing this method authors from  paper were able to improve Imagenet score from</code> 86.4 <code>to</code>88.4`</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F8c006f7e5e33cebbd7a72c135fa0b6f2%2FScreen%20Shot%202020-01-28%20at%2011.27.49%20AM.png?generation=1580228955932575&amp;alt=media\" alt=\"\"></p>\n\n<p>I made a simple kernel that give you example how we can generate random Bengali data which can be used for generating pseudo labels. Obviously this kernel can be improved we can also add more publicly available data. </p>\n\n<p>link:  <a href=\"https://www.kaggle.com/drhabib/generating-more-training-data\">https://www.kaggle.com/drhabib/generating-more-training-data</a></p>\n\n<p>Example Images:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F2f5f1760b623d078ef30ad0883b141fb%2FScreen%20Shot%202020-01-28%20at%2011.20.38%20AM.png?generation=1580228534527305&amp;alt=media\" alt=\"\"></p>\n\n<p>P.S my apologize if someone already posted this article and I failed to mention name =) </p>\n\n<p>Val An thanks for the discussion <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808\">https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808</a></p>",
      "rawMarkdown": "Hi everyone. \nRecently I came across this paper https://arxiv.org/pdf/1911.04252v2.pdf (please read it for details). Currently its hold record for top score on Imagenet classification \n\nTL DR from this paper::\n```\n1. Train a small network on your training data\n2) Use completely unrelated Images to make predictions and generate Pseudo Labels\n3) Train a bigger student network with your data, and Pseudo Labels \n\n```\nUsing this method authors from  paper were able to improve Imagenet score from` 86.4 `to `88.4`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F8c006f7e5e33cebbd7a72c135fa0b6f2%2FScreen%20Shot%202020-01-28%20at%2011.27.49%20AM.png?generation=1580228955932575&amp;alt=media)\n\n\nI made a simple kernel that give you example how we can generate random Bengali data which can be used for generating pseudo labels. Obviously this kernel can be improved we can also add more publicly available data. \n\nlink:  https://www.kaggle.com/drhabib/generating-more-training-data\n\nExample Images:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F2f5f1760b623d078ef30ad0883b141fb%2FScreen%20Shot%202020-01-28%20at%2011.20.38%20AM.png?generation=1580228534527305&amp;alt=media)\n\nP.S my apologize if someone already posted this article and I failed to mention name =) \n\nVal An thanks for the discussion https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808",
      "votes": null
    },
    {
      "id": "731477",
      "postDate": "01/28/2020 17:20:08",
      "content": "<p>Great approach. But rather than using a dictionary like this, would it be more promising to use existing Bangla character datasets to produce Pseudo Labels? I didn't try Pseudo Labels yet, just re-thinking to see your idea. </p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/122604\">https://www.kaggle.com/c/bengaliai-cv19/discussion/122604</a></p>",
      "rawMarkdown": "Great approach. But rather than using a dictionary like this, would it be more promising to use existing Bangla character datasets to produce Pseudo Labels? I didn't try Pseudo Labels yet, just re-thinking to see your idea. \n\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/122604",
      "votes": null
    },
    {
      "id": "731478",
      "postDate": "01/28/2020 17:22:45",
      "content": "<p>Mentioned here :) <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808\">https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808</a></p>",
      "rawMarkdown": "Mentioned here :) [https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808](https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808)",
      "votes": null
    },
    {
      "id": "731486",
      "postDate": "01/28/2020 17:30:05",
      "content": "<p>thanks! I edited the post! </p>",
      "rawMarkdown": "thanks! I edited the post!",
      "votes": null
    },
    {
      "id": "731496",
      "postDate": "01/28/2020 17:44:23",
      "content": "<p>I think anything remotely related can work =) We just have to try =) </p>",
      "rawMarkdown": "I think anything remotely related can work =) We just have to try =)",
      "votes": null
    },
    {
      "id": "731598",
      "postDate": "01/28/2020 20:04:15",
      "content": "<p>They also have <a href=\"https://github.com/MinhasKamal/BengaliDictionary/blob/master/BengaliCharacterCombinations.txt\">character combinations</a>. Maybe a GAN can be trained to construct missing combinations handwriting dataset.</p>\n\n<p>Alternatively, we can force a model to generalize well on unseen combinations with a validation strategy <a href=\"https://www.kaggle.com/sibmike/bengali-groupshufflesplit-validation\">described here</a></p>",
      "rawMarkdown": "They also have [character combinations](https://github.com/MinhasKamal/BengaliDictionary/blob/master/BengaliCharacterCombinations.txt). Maybe a GAN can be trained to construct missing combinations handwriting dataset.\n\nAlternatively, we can force a model to generalize well on unseen combinations with a validation strategy [described here](https://www.kaggle.com/sibmike/bengali-groupshufflesplit-validation)",
      "votes": null
    },
    {
      "id": "734667",
      "postDate": "02/01/2020 20:08:42",
      "content": "<p>hey <a href=\"/drhabib\">@drhabib</a>  what do you think about generating more data or balancing using the GAN approach? </p>\n\n<p>paper: <a href=\"https://arxiv.org/pdf/1803.09655.pdf\">https://arxiv.org/pdf/1803.09655.pdf</a>\ncode: <a href=\"https://github.com/IBM/BAGAN\">https://github.com/IBM/BAGAN</a></p>",
      "rawMarkdown": "hey @drhabib  what do you think about generating more data or balancing using the GAN approach? \n\npaper: https://arxiv.org/pdf/1803.09655.pdf\ncode: https://github.com/IBM/BAGAN",
      "votes": null
    },
    {
      "id": "737155",
      "postDate": "02/05/2020 01:29:26",
      "content": "<p>Newer come to kaggle.I don't understand the Kaggle rules about External Data.what does \"post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\" mean? if i generating More data ,should i send a e-mail to kaggle?</p>",
      "rawMarkdown": "Newer come to kaggle.I don't understand the Kaggle rules about External Data.what does \"post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\" mean? if i generating More data ,should i send a e-mail to kaggle?",
      "votes": null
    },
    {
      "id": "737311",
      "postDate": "02/05/2020 07:07:52",
      "content": "<p>They mean any official external dataset such as some other language image data that's not from this competition, or model weights that you use to pretrain on that is not yours e.g. ImageNet weights as simple example. I think self-generated are fine.</p>",
      "rawMarkdown": "They mean any official external dataset such as some other language image data that's not from this competition, or model weights that you use to pretrain on that is not yours e.g. ImageNet weights as simple example. I think self-generated are fine.",
      "votes": null
    },
    {
      "id": "746605",
      "postDate": "02/15/2020 09:06:29",
      "content": "<p>Thank you for answer😄 </p>",
      "rawMarkdown": "Thank you for answer😄",
      "votes": null
    },
    {
      "id": "759892",
      "postDate": "02/29/2020 15:33:11",
      "content": "<p>Here's some results from my GAN approach, pretty useless thus far im afraid but it does look nice. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1110596%2F5cc312b9ae5f6f269bb036da8918278b%2FROOTS_CROP.jpg?generation=1582991217568198&amp;alt=media\" alt=\"]![\"></p>",
      "rawMarkdown": "Here's some results from my GAN approach, pretty useless thus far im afraid but it does look nice. ![]![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1110596%2F5cc312b9ae5f6f269bb036da8918278b%2FROOTS_CROP.jpg?generation=1582991217568198&amp;alt=media)",
      "votes": null
    },
    {
      "id": "763200",
      "postDate": "03/04/2020 08:27:02",
      "content": "<p>hey, did it help?</p>",
      "rawMarkdown": "hey, did it help?",
      "votes": null
    },
    {
      "id": "764721",
      "postDate": "03/05/2020 19:53:39",
      "content": "<p>I havn't had the time to test it out yet, I had issues creating a good dataset from GAN as characters shifts into one another if you vary \"style\" far from the mean. I hope to have more time experimenting before the comp ends. </p>",
      "rawMarkdown": "I havn't had the time to test it out yet, I had issues creating a good dataset from GAN as characters shifts into one another if you vary \"style\" far from the mean. I hope to have more time experimenting before the comp ends.",
      "votes": null
    },
    {
      "id": "764796",
      "postDate": "03/05/2020 22:56:11",
      "content": "<p>see if you can do a supervised version of gan:</p>\n\n<ol>\n<li><p>assume you have x1,x2,x3 ....xN images of the same class</p></li>\n<li><p>you want to find GAN(xi, unknown latent) = xj_bar , i.e. image xi can be perturbed/transformed so that it would look like another image xj  in the same class.</p></li>\n<li><p>besides the usual generator and discriminator adversarial loss, we also want min (least_sq(xj_bar-xk)) over all k in training set . other metric distance , either over image or feature, can also be used. other \"pooling\" can also be use, e.g. replace min to average, etc</p></li>\n<li><p>if it is successful, then you have a way to map from one image to another and you can use GAN as an augmentation transform</p></li>\n<li><p>idea can be extended to map from one class to another, etc</p></li>\n</ol>",
      "rawMarkdown": "see if you can do a supervised version of gan:\n\n1. assume you have x1,x2,x3 ....xN images of the same class\n\n2. you want to find GAN(xi, unknown latent) = xj\\_bar , i.e. image xi can be perturbed/transformed so that it would look like another image xj  in the same class.\n\n3. besides the usual generator and discriminator adversarial loss, we also want min (least\\_sq(xj\\_bar-xk)) over all k in training set . other metric distance , either over image or feature, can also be used. other \"pooling\" can also be use, e.g. replace min to average, etc\n\n4. if it is successful, then you have a way to map from one image to another and you can use GAN as an augmentation transform\n\n5. idea can be extended to map from one class to another, etc",
      "votes": null
    },
    {
      "id": "764807",
      "postDate": "03/05/2020 23:31:07",
      "content": "<p>there is yet another way to use the GAN data as unlablled training set via mixup.</p>\n\n<ol>\n<li>original mixup: \n<ul><li>x = a * xi + (1-a) * xj</li>\n<li>loss = a * loss(pi,truthi) + (1-a) * loss(pj,truthj)</li></ul></li>\n</ol>\n\n<p>since now the GAN data are unlabeled, one may use:</p>\n\n<ol>\n<li>unlabeled mixup (assume thruthj  is unknown): \n<ul><li>x = a * xi + (1-a) * xj</li>\n<li>loss = a * loss(pi,truthi)  </li></ul></li>\n</ol>",
      "rawMarkdown": "there is yet another way to use the GAN data as unlablled training set via mixup.\n\n1. original mixup: \n   - x = a * xi + (1-a) * xj\n   - loss = a * loss(pi,truthi) + (1-a) * loss(pj,truthj)\n\nsince now the GAN data are unlabeled, one may use:\n\n\n1. unlabeled mixup (assume thruthj  is unknown): \n   - x = a * xi + (1-a) * xj\n   - loss = a * loss(pi,truthi)",
      "votes": null
    },
    {
      "id": "764878",
      "postDate": "03/06/2020 02:22:49",
      "content": "<p>Nice job Max. Those GAN images look cool.</p>",
      "rawMarkdown": "Nice job Max. Those GAN images look cool.",
      "votes": null
    },
    {
      "id": "764962",
      "postDate": "03/06/2020 06:10:39",
      "content": "<p>Check this new paper today\n<a href=\"https://arxiv.org/pdf/2003.02567.pdf\">https://arxiv.org/pdf/2003.02567.pdf</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F7e06301ca25a724f1b11636aa1fda8c4%2F35B046D7-2818-4CBF-ADBF-D85A875BC288.png?generation=1583475035735033&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Check this new paper today\nhttps://arxiv.org/pdf/2003.02567.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F7e06301ca25a724f1b11636aa1fda8c4%2F35B046D7-2818-4CBF-ADBF-D85A875BC288.png?generation=1583475035735033&amp;alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 731477,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "01/28/2020 17:20:08",
      "content": "<p>Great approach. But rather than using a dictionary like this, would it be more promising to use existing Bangla character datasets to produce Pseudo Labels? I didn't try Pseudo Labels yet, just re-thinking to see your idea. </p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/122604\">https://www.kaggle.com/c/bengaliai-cv19/discussion/122604</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 731496,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "01/28/2020 17:44:23",
          "content": "<p>I think anything remotely related can work =) We just have to try =) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 734667,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "02/01/2020 20:08:42",
          "content": "<p>hey <a href=\"/drhabib\">@drhabib</a>  what do you think about generating more data or balancing using the GAN approach? </p>\n\n<p>paper: <a href=\"https://arxiv.org/pdf/1803.09655.pdf\">https://arxiv.org/pdf/1803.09655.pdf</a>\ncode: <a href=\"https://github.com/IBM/BAGAN\">https://github.com/IBM/BAGAN</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 731478,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "01/28/2020 17:22:45",
      "content": "<p>Mentioned here :) <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808\">https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 731486,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "01/28/2020 17:30:05",
          "content": "<p>thanks! I edited the post! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 731598,
      "author_name": "sibmike",
      "author_url": "",
      "post_date": "01/28/2020 20:04:15",
      "content": "<p>They also have <a href=\"https://github.com/MinhasKamal/BengaliDictionary/blob/master/BengaliCharacterCombinations.txt\">character combinations</a>. Maybe a GAN can be trained to construct missing combinations handwriting dataset.</p>\n\n<p>Alternatively, we can force a model to generalize well on unseen combinations with a validation strategy <a href=\"https://www.kaggle.com/sibmike/bengali-groupshufflesplit-validation\">described here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 737155,
      "author_name": "raininbox",
      "author_url": "",
      "post_date": "02/05/2020 01:29:26",
      "content": "<p>Newer come to kaggle.I don't understand the Kaggle rules about External Data.what does \"post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\" mean? if i generating More data ,should i send a e-mail to kaggle?</p>",
      "votes": null,
      "replies": [
        {
          "id": 737311,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "02/05/2020 07:07:52",
          "content": "<p>They mean any official external dataset such as some other language image data that's not from this competition, or model weights that you use to pretrain on that is not yours e.g. ImageNet weights as simple example. I think self-generated are fine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746605,
          "author_name": "raininbox",
          "author_url": "",
          "post_date": "02/15/2020 09:06:29",
          "content": "<p>Thank you for answer😄 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 759892,
      "author_name": "maxjon",
      "author_url": "",
      "post_date": "02/29/2020 15:33:11",
      "content": "<p>Here's some results from my GAN approach, pretty useless thus far im afraid but it does look nice. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1110596%2F5cc312b9ae5f6f269bb036da8918278b%2FROOTS_CROP.jpg?generation=1582991217568198&amp;alt=media\" alt=\"]![\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 763200,
          "author_name": "ankitsajwan",
          "author_url": "",
          "post_date": "03/04/2020 08:27:02",
          "content": "<p>hey, did it help?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764721,
          "author_name": "maxjon",
          "author_url": "",
          "post_date": "03/05/2020 19:53:39",
          "content": "<p>I havn't had the time to test it out yet, I had issues creating a good dataset from GAN as characters shifts into one another if you vary \"style\" far from the mean. I hope to have more time experimenting before the comp ends. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764796,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/05/2020 22:56:11",
          "content": "<p>see if you can do a supervised version of gan:</p>\n\n<ol>\n<li><p>assume you have x1,x2,x3 ....xN images of the same class</p></li>\n<li><p>you want to find GAN(xi, unknown latent) = xj_bar , i.e. image xi can be perturbed/transformed so that it would look like another image xj  in the same class.</p></li>\n<li><p>besides the usual generator and discriminator adversarial loss, we also want min (least_sq(xj_bar-xk)) over all k in training set . other metric distance , either over image or feature, can also be used. other \"pooling\" can also be use, e.g. replace min to average, etc</p></li>\n<li><p>if it is successful, then you have a way to map from one image to another and you can use GAN as an augmentation transform</p></li>\n<li><p>idea can be extended to map from one class to another, etc</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764807,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/05/2020 23:31:07",
          "content": "<p>there is yet another way to use the GAN data as unlablled training set via mixup.</p>\n\n<ol>\n<li>original mixup: \n<ul><li>x = a * xi + (1-a) * xj</li>\n<li>loss = a * loss(pi,truthi) + (1-a) * loss(pj,truthj)</li></ul></li>\n</ol>\n\n<p>since now the GAN data are unlabeled, one may use:</p>\n\n<ol>\n<li>unlabeled mixup (assume thruthj  is unknown): \n<ul><li>x = a * xi + (1-a) * xj</li>\n<li>loss = a * loss(pi,truthi)  </li></ul></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764878,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/06/2020 02:22:49",
          "content": "<p>Nice job Max. Those GAN images look cool.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764962,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/06/2020 06:10:39",
          "content": "<p>Check this new paper today\n<a href=\"https://arxiv.org/pdf/2003.02567.pdf\">https://arxiv.org/pdf/2003.02567.pdf</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F7e06301ca25a724f1b11636aa1fda8c4%2F35B046D7-2818-4CBF-ADBF-D85A875BC288.png?generation=1583475035735033&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "731429": "Hi everyone. \nRecently I came across this paper https://arxiv.org/pdf/1911.04252v2.pdf (please read it for details). Currently its hold record for top score on Imagenet classification \n\nTL DR from this paper::\n```\n1. Train a small network on your training data\n2) Use completely unrelated Images to make predictions and generate Pseudo Labels\n3) Train a bigger student network with your data, and Pseudo Labels \n\n```\nUsing this method authors from  paper were able to improve Imagenet score from` 86.4 `to `88.4`\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F8c006f7e5e33cebbd7a72c135fa0b6f2%2FScreen%20Shot%202020-01-28%20at%2011.27.49%20AM.png?generation=1580228955932575&amp;alt=media)\n\n\nI made a simple kernel that give you example how we can generate random Bengali data which can be used for generating pseudo labels. Obviously this kernel can be improved we can also add more publicly available data. \n\nlink:  https://www.kaggle.com/drhabib/generating-more-training-data\n\nExample Images:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F2f5f1760b623d078ef30ad0883b141fb%2FScreen%20Shot%202020-01-28%20at%2011.20.38%20AM.png?generation=1580228534527305&amp;alt=media)\n\nP.S my apologize if someone already posted this article and I failed to mention name =) \n\nVal An thanks for the discussion https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808",
    "731477": "Great approach. But rather than using a dictionary like this, would it be more promising to use existing Bangla character datasets to produce Pseudo Labels? I didn't try Pseudo Labels yet, just re-thinking to see your idea. \n\nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/122604",
    "731478": "Mentioned here :) [https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808](https://www.kaggle.com/c/bengaliai-cv19/discussion/125046#713808)",
    "731486": "thanks! I edited the post!",
    "731496": "I think anything remotely related can work =) We just have to try =)",
    "731598": "They also have [character combinations](https://github.com/MinhasKamal/BengaliDictionary/blob/master/BengaliCharacterCombinations.txt). Maybe a GAN can be trained to construct missing combinations handwriting dataset.\n\nAlternatively, we can force a model to generalize well on unseen combinations with a validation strategy [described here](https://www.kaggle.com/sibmike/bengali-groupshufflesplit-validation)",
    "734667": "hey @drhabib  what do you think about generating more data or balancing using the GAN approach? \n\npaper: https://arxiv.org/pdf/1803.09655.pdf\ncode: https://github.com/IBM/BAGAN",
    "737155": "Newer come to kaggle.I don't understand the Kaggle rules about External Data.what does \"post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\" mean? if i generating More data ,should i send a e-mail to kaggle?",
    "737311": "They mean any official external dataset such as some other language image data that's not from this competition, or model weights that you use to pretrain on that is not yours e.g. ImageNet weights as simple example. I think self-generated are fine.",
    "746605": "Thank you for answer😄",
    "759892": "Here's some results from my GAN approach, pretty useless thus far im afraid but it does look nice. ![]![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1110596%2F5cc312b9ae5f6f269bb036da8918278b%2FROOTS_CROP.jpg?generation=1582991217568198&amp;alt=media)",
    "763200": "hey, did it help?",
    "764721": "I havn't had the time to test it out yet, I had issues creating a good dataset from GAN as characters shifts into one another if you vary \"style\" far from the mean. I hope to have more time experimenting before the comp ends.",
    "764796": "see if you can do a supervised version of gan:\n\n1. assume you have x1,x2,x3 ....xN images of the same class\n\n2. you want to find GAN(xi, unknown latent) = xj\\_bar , i.e. image xi can be perturbed/transformed so that it would look like another image xj  in the same class.\n\n3. besides the usual generator and discriminator adversarial loss, we also want min (least\\_sq(xj\\_bar-xk)) over all k in training set . other metric distance , either over image or feature, can also be used. other \"pooling\" can also be use, e.g. replace min to average, etc\n\n4. if it is successful, then you have a way to map from one image to another and you can use GAN as an augmentation transform\n\n5. idea can be extended to map from one class to another, etc",
    "764807": "there is yet another way to use the GAN data as unlablled training set via mixup.\n\n1. original mixup: \n   - x = a * xi + (1-a) * xj\n   - loss = a * loss(pi,truthi) + (1-a) * loss(pj,truthj)\n\nsince now the GAN data are unlabeled, one may use:\n\n\n1. unlabeled mixup (assume thruthj  is unknown): \n   - x = a * xi + (1-a) * xj\n   - loss = a * loss(pi,truthi)",
    "764878": "Nice job Max. Those GAN images look cool.",
    "764962": "Check this new paper today\nhttps://arxiv.org/pdf/2003.02567.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F7e06301ca25a724f1b11636aa1fda8c4%2F35B046D7-2818-4CBF-ADBF-D85A875BC288.png?generation=1583475035735033&amp;alt=media)"
  },
  "source": "meta"
}