{
  "id": 123757,
  "title": "[placeholder] lb0.963+ pytorch starter kit",
  "url": "/competitions/bengaliai-cv19/discussion/123757",
  "author_name": "hengck23",
  "post_date": "2019-12-30T07:13:24.061000",
  "votes": 173,
  "comment_count": 123,
  "views": 0,
  "content": "<p>version 20191230:\n- densenet 121 ensemble + swa without cyclic rate, without bn refinement\n- inference at 25 min</p>\n\n<hr>\n\n<p>version 20200111:\n- se-resnext50  + balanced sampler + mixup (please refer to readme.ppt at the google drive</p>\n\n<hr>\n\n<p>version 20200224:\n- modified se-resnext50 with augmentation that works. please see read file in folder. This version can gives CV 0.991 and LB 0.980  </p>\n\n<hr>\n\n<p><a href=\"https://drive.google.com/open?id=1A5CygainZ4rO_rOs4mUjQk378WlF8MrN\">https://drive.google.com/open?id=1A5CygainZ4rO_rOs4mUjQk378WlF8MrN</a></p>\n\n<p>... to be updated ...\ne.g. \n- next version mixup,cutout, manifold mixup, adversarial loss?\nsee <a href=\"https://github.com/rois-codh/kmnist\">https://github.com/rois-codh/kmnist</a></p>\n\n<ul>\n<li><p>metric learning loss ... large-margin softmax, cosine, center loss?</p></li>\n<li><p>long tail, distribution, loss weighing, balance sampling, etc?</p></li>\n<li><p>sub class, metric distance without triplet</p></li>\n<li><p>attention-based (stroke based?), human in-the-loop</p></li>\n<li><p>LSTM, RNN, CTC loss, stroke parsing?</p></li>\n<li><p>spatial transformation net, bounding box normalisation </p></li>\n<li>augmentation, add random  distractor strokes</li>\n<li>treating as segmentation problem, or pixel classification + novel pooling</li>\n<li><p>feature embedding (<a href=\"https://github.com/qychen13/DifficultyAwareEmbedding\">https://github.com/qychen13/DifficultyAwareEmbedding</a>), metric learning (<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109</a>)</p></li>\n<li><p>fine grained, attribute learning</p></li>\n<li>multi-task and task dependency</li>\n</ul>",
  "messages": [
    {
      "id": 706279,
      "postDate": "2019-12-30T07:13:24.060Z",
      "content": "<p>version 20191230:\n- densenet 121 ensemble + swa without cyclic rate, without bn refinement\n- inference at 25 min</p>\n\n<hr>\n\n<p>version 20200111:\n- se-resnext50  + balanced sampler + mixup (please refer to readme.ppt at the google drive</p>\n\n<hr>\n\n<p>version 20200224:\n- modified se-resnext50 with augmentation that works. please see read file in folder. This version can gives CV 0.991 and LB 0.980  </p>\n\n<hr>\n\n<p><a href=\"https://drive.google.com/open?id=1A5CygainZ4rO_rOs4mUjQk378WlF8MrN\">https://drive.google.com/open?id=1A5CygainZ4rO_rOs4mUjQk378WlF8MrN</a></p>\n\n<p>... to be updated ...\ne.g. \n- next version mixup,cutout, manifold mixup, adversarial loss?\nsee <a href=\"https://github.com/rois-codh/kmnist\">https://github.com/rois-codh/kmnist</a></p>\n\n<ul>\n<li><p>metric learning loss ... large-margin softmax, cosine, center loss?</p></li>\n<li><p>long tail, distribution, loss weighing, balance sampling, etc?</p></li>\n<li><p>sub class, metric distance without triplet</p></li>\n<li><p>attention-based (stroke based?), human in-the-loop</p></li>\n<li><p>LSTM, RNN, CTC loss, stroke parsing?</p></li>\n<li><p>spatial transformation net, bounding box normalisation </p></li>\n<li>augmentation, add random  distractor strokes</li>\n<li>treating as segmentation problem, or pixel classification + novel pooling</li>\n<li><p>feature embedding (<a href=\"https://github.com/qychen13/DifficultyAwareEmbedding\">https://github.com/qychen13/DifficultyAwareEmbedding</a>), metric learning (<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109</a>)</p></li>\n<li><p>fine grained, attribute learning</p></li>\n<li>multi-task and task dependency</li>\n</ul>",
      "rawMarkdown": "version 20191230:\n- densenet 121 ensemble + swa without cyclic rate, without bn refinement\n- inference at 25 min\n\n---\n\nversion 20200111:\n- se-resnext50  + balanced sampler + mixup (please refer to readme.ppt at the google drive\n \n\n\n---\n\nversion 20200224:\n- modified se-resnext50 with augmentation that works. please see read file in folder. This version can gives CV 0.991 and LB 0.980  \n \n\n---\n\n\n\nhttps://drive.google.com/open?id=1A5CygainZ4rO_rOs4mUjQk378WlF8MrN\n\n... to be updated ...\ne.g. \n- next version mixup,cutout, manifold mixup, adversarial loss?\nsee https://github.com/rois-codh/kmnist\n       \n- metric learning loss ... large-margin softmax, cosine, center loss?\n\n- long tail, distribution, loss weighing, balance sampling, etc?\n\n- sub class, metric distance without triplet\n\n- attention-based (stroke based?), human in-the-loop\n\n- LSTM, RNN, CTC loss, stroke parsing?\n\n- spatial transformation net, bounding box normalisation \n- augmentation, add random  distractor strokes\n- treating as segmentation problem, or pixel classification + novel pooling\n- feature embedding (https://github.com/qychen13/DifficultyAwareEmbedding), metric learning (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109)\n\n- fine grained, attribute learning\n- multi-task and task dependency",
      "votes": 173
    },
    {
      "id": 771046,
      "postDate": "2020-03-13T17:42:27.323Z",
      "content": "<p>this is one of the kaggle competitions that i did the most number of experiments. I notice something interesting.</p>\n\n<p>for several experiments, it is important to apply the method at the early start of training. e.g. you get different results if you apply ohem on trained models (i.e. end stage of training) versus you apply it right at the start of training.</p>\n\n<p>i think that when data size is small, the number of good solutions (or generalized solutions) region is small. i.e. there are many local minimum. It is difficult to jump from one solution to another solution </p>",
      "rawMarkdown": "this is one of the kaggle competitions that i did the most number of experiments. I notice something interesting.\n\nfor several experiments, it is important to apply the method at the early start of training. e.g. you get different results if you apply ohem on trained models (i.e. end stage of training) versus you apply it right at the start of training.\n\ni think that when data size is small, the number of good solutions (or generalized solutions) region is small. i.e. there are many local minimum. It is difficult to jump from one solution to another solution \n",
      "votes": 8,
      "replies": [
        {
          "id": 771347,
          "postDate": "2020-03-14T04:07:33.507Z",
          "content": "<p>I do find out applying ohem at the end stage of training will save you from overfitting. It will shake a little bit at the beginning but it will increase your local cv at the end. It seems like applying ohem make it jumping out of its local minimum.</p>",
          "rawMarkdown": "I do find out applying ohem at the end stage of training will save you from overfitting. It will shake a little bit at the beginning but it will increase your local cv at the end. It seems like applying ohem make it jumping out of its local minimum.",
          "votes": 3
        }
      ]
    },
    {
      "id": 756402,
      "postDate": "2020-02-25T17:37:08.607Z",
      "content": "<p>this is a reference  version for those who are struggling:</p>\n\n<p>version 20200224:\n- modified se-resnext50 (64x112 small input) with augmentation that works. please see readme file in folder. \n  This version can gives CV 0.991 and LB 0.980 after you optimized the augmentation hyper parameters  yourselves.</p>\n\n<ul>\n<li><p>it provides reference log file for you to compare the loss curve</p></li>\n<li><p>it has train/validation split that has gap of CV/LB 0.011 </p></li>\n<li><p>code is not complete, but should have enough details to reproduce the above mentioned results.\n(do not request me for missing files or functions)</p></li>\n</ul>",
      "rawMarkdown": "this is a reference  version for those who are struggling:\n\nversion 20200224:\n- modified se-resnext50 (64x112 small input) with augmentation that works. please see readme file in folder. \n  This version can gives CV 0.991 and LB 0.980 after you optimized the augmentation hyper parameters  yourselves.\n \n- it provides reference log file for you to compare the loss curve\n\n- it has train/validation split that has gap of CV/LB 0.011 \n\n- code is not complete, but should have enough details to reproduce the above mentioned results.\n  (do not request me for missing files or functions)",
      "votes": 7,
      "replies": [
        {
          "id": 756544,
          "postDate": "2020-02-25T20:23:17.407Z",
          "content": "<p>Thanks for sharing and for the best thread in the discussion!</p>",
          "rawMarkdown": "Thanks for sharing and for the best thread in the discussion!"
        },
        {
          "id": 756562,
          "postDate": "2020-02-25T20:56:28.917Z",
          "content": "<p>Hi Heng Thanks for the update. May I ask if the 64x112 is a simple downsize from the original 137x236 image or did you do some post processing?</p>\n\n<p>And may I ask how long have you trained the model to reach cv ~.99. From that log file it seems it's been training for 500 epochs?!</p>",
          "rawMarkdown": "Hi Heng Thanks for the update. May I ask if the 64x112 is a simple downsize from the original 137x236 image or did you do some post processing?\n\nAnd may I ask how long have you trained the model to reach cv ~.99. From that log file it seems it's been training for 500 epochs?!"
        },
        {
          "id": 756712,
          "postDate": "2020-02-26T01:43:17.457Z",
          "content": "<p>please refer to the readme file</p>",
          "rawMarkdown": "please refer to the readme file",
          "votes": 2
        },
        {
          "id": 757358,
          "postDate": "2020-02-26T17:09:16.117Z",
          "content": "<p>how to train a modified resenext model with pretrained imagenet model</p>\n\n<ol>\n<li><p>use the unmodified version resenext model first. initialised with pretrained imagenet model and train as usual.</p></li>\n<li><p>assume resenext = block0, block1 ... block4. we modify block0.</p></li>\n<li><p>load the previously trained model except for the the modified part (i.e block0).</p></li>\n<li><p>freeze all modified part (including the classifier head). Train and the gradient will only back-propagated at the modified block0.</p></li>\n<li><p>when the accuracy/loss for the modified model is about the same as the previously unmodified one, say 90% of the results, unfreeze all  model and proceed training as usual.</p></li>\n</ol>\n\n<p>this is how i do network surgery. apply this when there is structural change of the model, e.g. add new convolution layers, reduce number of channels, change kernel size, ...</p>\n\n<hr>\n\n<p>if the old pretrained imagenet model can be loaded after modification (e.g. simple modification like changing strike), there is no need to freeze and unfreeze. In such minor modification, just train as usual.</p>",
          "rawMarkdown": "how to train a modified resenext model with pretrained imagenet model\n\n1. use the unmodified version resenext model first. initialised with pretrained imagenet model and train as usual.\n\n2. assume resenext = block0, block1 ... block4. we modify block0.\n\n3. load the previously trained model except for the the modified part (i.e block0).\n\n4. freeze all modified part (including the classifier head). Train and the gradient will only back-propagated at the modified block0.\n\n5. when the accuracy/loss for the modified model is about the same as the previously unmodified one, say 90% of the results, unfreeze all  model and proceed training as usual.\n\nthis is how i do network surgery. apply this when there is structural change of the model, e.g. add new convolution layers, reduce number of channels, change kernel size, ...\n\n---\n\nif the old pretrained imagenet model can be loaded after modification (e.g. simple modification like changing strike), there is no need to freeze and unfreeze. In such minor modification, just train as usual.\n ",
          "votes": 3
        },
        {
          "id": 758076,
          "postDate": "2020-02-27T12:16:27.667Z",
          "content": "<p>purely trained from scratch \"without surgery above\".\nno augmentation is used at all!</p>\n\n<p>local cv = 0.972 (at rate 0.05)\nlocal cv = 0.972 (at rate 0.005)</p>\n\n<p>training loss is almost zero, so there is no point to train further. expected LB should be around 0.960 (without augmentation)</p>\n\n<p>see attached file.\n```\ndef valid_augment(image, label, infor):\n    image = to_64x112(image)\n    return image, label, infor</p>\n\n<p>def train_augment(image, label, infor):\n    original = image.copy()\n    original = to_64x112(original)\n    image = to_64x112(image)\n    return original, image, label, infor</p>\n\n<p>def train_batch_augment(original, input, onehot):\n    return input, onehot</p>\n\n<p>0.05000   9.0   15.6 | 0.972 : 0.957 0.989 0.982  0.955 | 0.20, 0.06, 0.06, 0.210 : 0.96, 0.99, 0.99, 0.956 | 0.00, 0.00, 0.00, 0.009 | 1 hr 38 min\n0.05000  10.0   17.3 | 0.972 : 0.957 0.990 0.986  0.955 | 0.21, 0.07, 0.07, 0.217 : 0.96, 0.99, 0.99, 0.956 | 0.00, 0.00, 0.00, 0.007 | 1 hr 49 min\n0.05000  11.0   19.0 | 0.971 : 0.957 0.989 0.982  0.955 | 0.21, 0.06, 0.07, 0.214 : 0.96, 0.99, 0.99, 0.956 | 0.01, 0.00, 0.00, 0.008 | 2 hr 00 min\n0.05000  12.0*  20.8 | 0.970 : 0.955 0.990 0.982  0.953 | 0.22, 0.06, 0.07, 0.224 : 0.96, 0.99, 0.99, 0.954 | 0.00, 0.00, 0.00, 0.007 | 2 hr 11 min\n```\nnote: some kaggler reported about to get LB 0.974 without augmentation (which i think 0.984 for local CV). hence i need to investigate better model or regularization (like dropout, shakedrop, etc) or better optimiser</p>",
          "rawMarkdown": "purely trained from scratch \"without surgery above\".\nno augmentation is used at all!\n\nlocal cv = 0.972 (at rate 0.05)\nlocal cv = 0.972 (at rate 0.005)\n\ntraining loss is almost zero, so there is no point to train further. expected LB should be around 0.960 (without augmentation)\n\n\nsee attached file.\n```\ndef valid_augment(image, label, infor):\n    image = to_64x112(image)\n    return image, label, infor\n\n\ndef train_augment(image, label, infor):\n    original = image.copy()\n    original = to_64x112(original)\n    image = to_64x112(image)\n    return original, image, label, infor\n\ndef train_batch_augment(original, input, onehot):\n    return input, onehot\n\n0.05000   9.0   15.6 | 0.972 : 0.957 0.989 0.982  0.955 | 0.20, 0.06, 0.06, 0.210 : 0.96, 0.99, 0.99, 0.956 | 0.00, 0.00, 0.00, 0.009 | 1 hr 38 min\n0.05000  10.0   17.3 | 0.972 : 0.957 0.990 0.986  0.955 | 0.21, 0.07, 0.07, 0.217 : 0.96, 0.99, 0.99, 0.956 | 0.00, 0.00, 0.00, 0.007 | 1 hr 49 min\n0.05000  11.0   19.0 | 0.971 : 0.957 0.989 0.982  0.955 | 0.21, 0.06, 0.07, 0.214 : 0.96, 0.99, 0.99, 0.956 | 0.01, 0.00, 0.00, 0.008 | 2 hr 00 min\n0.05000  12.0*  20.8 | 0.970 : 0.955 0.990 0.982  0.953 | 0.22, 0.06, 0.07, 0.224 : 0.96, 0.99, 0.99, 0.954 | 0.00, 0.00, 0.00, 0.007 | 2 hr 11 min\n```\nnote: some kaggler reported about to get LB 0.974 without augmentation (which i think 0.984 for local CV). hence i need to investigate better model or regularization (like dropout, shakedrop, etc) or better optimiser",
          "votes": 2
        },
        {
          "id": 758342,
          "postDate": "2020-02-27T16:49:19.633Z",
          "content": "<p>additional information:</p>\n\n<p>modified se-resnext50 (32x56 mini input) with same augmentation that works  gives CV 0.987\nthis may be useful for auto-augmentation parameters search?</p>",
          "rawMarkdown": "additional information:\n\nmodified se-resnext50 (32x56 mini input) with same augmentation that works  gives CV 0.987\nthis may be useful for auto-augmentation parameters search?\n\n \n",
          "votes": 4
        },
        {
          "id": 759575,
          "postDate": "2020-02-29T07:58:51.023Z",
          "content": "<p>Thanks for your sharing,but I have a question.\nIn your train_rand_aug2c.py,there is a func named INVERSE_COMPOSE,but you didn't assign it before using，I guess it is used to add information for first three predict form firth .\nCould you explain its meaning in detail? Thanks！</p>",
          "rawMarkdown": "Thanks for your sharing,but I have a question.\nIn your train_rand_aug2c.py,there is a func named INVERSE_COMPOSE,but you didn't assign it before using，I guess it is used to add information for first three predict form firth .\nCould you explain its meaning in detail? Thanks！"
        },
        {
          "id": 759588,
          "postDate": "2020-02-29T08:24:02.713Z",
          "content": "<p>we know that each of the 1295 grapheme is made up of grapheme=root, vowel, consonant</p>\n\n<p>so if you have predicted the 1295 class, simply  decode (root, vowel, consonant) = grapheme</p>",
          "rawMarkdown": "we know that each of the 1295 grapheme is made up of grapheme=root, vowel, consonant\n\n\nso if you have predicted the 1295 class, simply  decode (root, vowel, consonant) = grapheme"
        },
        {
          "id": 759960,
          "postDate": "2020-02-29T16:56:38.817Z",
          "content": "<p>Oh,yes.I understand your idea.</p>\n\n<p>But I got another question when I use your code to train the model .Most of the time my GPU utilization ratio is 0 and I thought the reason for this is dataloder's speed limitation.When we load the data by dataloader,we do some augmentation,and before send data to the net,maybe do some mixup or cutout.These process are not handled by GPU but CPU.</p>\n\n<p>This problem made my training process very slow,but in your log.train,it shows the speed is not so slow,so I want to ask whether you use some tricks to speed up.</p>",
          "rawMarkdown": "Oh,yes.I understand your idea.\n\nBut I got another question when I use your code to train the model .Most of the time my GPU utilization ratio is 0 and I thought the reason for this is dataloder's speed limitation.When we load the data by dataloader,we do some augmentation,and before send data to the net,maybe do some mixup or cutout.These process are not handled by GPU but CPU.\n\nThis problem made my training process very slow,but in your log.train,it shows the speed is not so slow,so I want to ask whether you use some tricks to speed up.\n"
        },
        {
          "id": 759969,
          "postDate": "2020-02-29T17:13:32.273Z",
          "content": "<p>df['xxx'].values should be done once at initialisation and not everytime at getitem()</p>",
          "rawMarkdown": "df['xxx'].values should be done once at initialisation and not everytime at getitem()"
        },
        {
          "id": 759981,
          "postDate": "2020-02-29T17:27:09.323Z",
          "content": "<p>```</p>\n\n<p>class KaggleDataset(Dataset):\n    def <strong>init</strong>(self, split, mode, csv, parquet, augment=None):\n        global TRAIN_PARQUET\n        ...\n        self.df_values = df.values</p>\n\n<pre><code>def __getitem__(self, index): \n    i, image_id, grapheme_root, vowel_diacritic, consonant_diacritic, grapheme, grapheme_symbol  =  self.df_values[index]\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "```\n\n\nclass KaggleDataset(Dataset):\n    def __init__(self, split, mode, csv, parquet, augment=None):\n        global TRAIN_PARQUET\n        ...\n        self.df_values = df.values\n\n    def __getitem__(self, index): \n        i, image_id, grapheme_root, vowel_diacritic, consonant_diacritic, grapheme, grapheme_symbol  =  self.df_values[index]\n\n```"
        },
        {
          "id": 760463,
          "postDate": "2020-03-01T10:33:43.803Z",
          "content": "<p>Thank you,I will have a try.</p>",
          "rawMarkdown": "Thank you,I will have a try."
        },
        {
          "id": 764147,
          "postDate": "2020-03-05T07:19:47Z",
          "content": "<p>Thanks for sharing, very helpful! Quick question, when I'm using your model <code>ResNext50</code>, it seems that the model takes more than 11GB of GPU memory. Did you encounter similar issue? </p>\n\n<p>Or does this model only work when you freeze part of the model and only train its head?</p>\n\n<p>Thanks. </p>",
          "rawMarkdown": "Thanks for sharing, very helpful! Quick question, when I'm using your model `ResNext50`, it seems that the model takes more than 11GB of GPU memory. Did you encounter similar issue? \n\nOr does this model only work when you freeze part of the model and only train its head?\n\nThanks. "
        },
        {
          "id": 765755,
          "postDate": "2020-03-07T03:43:19.483Z",
          "content": "<p>I am still trying to figure out how to write inverse_compose function.  Any idea ?</p>",
          "rawMarkdown": "I am still trying to figure out how to write inverse_compose function.  Any idea ?"
        },
        {
          "id": 766504,
          "postDate": "2020-03-08T09:09:05.520Z",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> </p>\n\n<p>```\ndef run_prepare_data4():\n    df_grapheme = pd.read_csv(DATA_DIR+'/grapheme_1295.csv')\n    df_root = pd.read_csv(DATA_DIR+'/root_168.csv')\n    df_vowel = pd.read_csv(DATA_DIR+'/vowel_11.csv')\n    df_consonant = pd.read_csv(DATA_DIR+'/consonant_7.csv')</p>\n\n<pre><code>grapheme_map = dict(df_grapheme[['grapheme','label']].values)\nroot_map = dict(df_root[['root','label']].values)\nvowel_map = dict(df_vowel[['vowel','label']].values)\nconsonant_map = dict(df_consonant[['consonant','label']].values)\n\ngrapheme_inverse_map = dict(df_grapheme[['label','grapheme']].values)\nroot_inverse_map = dict(df_root[['label','root']].values)\nvowel_inverse_map = dict(df_vowel[['label','vowel']].values)\nconsonant_inverse_map = dict(df_consonant[['label','consonant',]].values)\n\n#----\ndf_train = pd.read_csv(DATA_DIR+'/train.csv')\ndf_train['grapheme']=df_train['grapheme'].map(grapheme_map)\nassert(df_train['grapheme'].isnull().values.any() == False)\n\ncompose = {}\ninverse_compose = {}\ngb = df_train.groupby(['grapheme',])\nfor i in range(1295):\n    gp = gb.get_group(i) #.index\n\n    gp = gp.drop_duplicates(subset=['grapheme_root', 'vowel_diacritic', 'consonant_diacritic'])\n    assert(len(gp)==1)\n\n    g = i\n    r,v,c = gp[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']].values[0]\n\n    inverse_compose[g]= (r,v,c)\n    compose[grapheme_inverse_map[g]]= (\n        root_inverse_map[r],\n        vowel_inverse_map[v],\n        consonant_inverse_map[c]\n    )\n\nprint(inverse_compose)\nprint(compose)\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "@phoenix9032 \n\n```\ndef run_prepare_data4():\n    df_grapheme = pd.read_csv(DATA_DIR+'/grapheme_1295.csv')\n    df_root = pd.read_csv(DATA_DIR+'/root_168.csv')\n    df_vowel = pd.read_csv(DATA_DIR+'/vowel_11.csv')\n    df_consonant = pd.read_csv(DATA_DIR+'/consonant_7.csv')\n\n\n    grapheme_map = dict(df_grapheme[['grapheme','label']].values)\n    root_map = dict(df_root[['root','label']].values)\n    vowel_map = dict(df_vowel[['vowel','label']].values)\n    consonant_map = dict(df_consonant[['consonant','label']].values)\n\n    grapheme_inverse_map = dict(df_grapheme[['label','grapheme']].values)\n    root_inverse_map = dict(df_root[['label','root']].values)\n    vowel_inverse_map = dict(df_vowel[['label','vowel']].values)\n    consonant_inverse_map = dict(df_consonant[['label','consonant',]].values)\n\n    #----\n    df_train = pd.read_csv(DATA_DIR+'/train.csv')\n    df_train['grapheme']=df_train['grapheme'].map(grapheme_map)\n    assert(df_train['grapheme'].isnull().values.any() == False)\n\n    compose = {}\n    inverse_compose = {}\n    gb = df_train.groupby(['grapheme',])\n    for i in range(1295):\n        gp = gb.get_group(i) #.index\n\n        gp = gp.drop_duplicates(subset=['grapheme_root', 'vowel_diacritic', 'consonant_diacritic'])\n        assert(len(gp)==1)\n\n        g = i\n        r,v,c = gp[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']].values[0]\n\n        inverse_compose[g]= (r,v,c)\n        compose[grapheme_inverse_map[g]]= (\n            root_inverse_map[r],\n            vowel_inverse_map[v],\n            consonant_inverse_map[c]\n        )\n\n    print(inverse_compose)\n    print(compose)\n\n```",
          "votes": 1
        },
        {
          "id": 766512,
          "postDate": "2020-03-08T09:21:13.480Z",
          "content": "<p>Thanks so much Heng :) </p>",
          "rawMarkdown": "Thanks so much Heng :) "
        }
      ]
    },
    {
      "id": 737223,
      "postDate": "2020-02-05T04:15:29.610Z",
      "content": "<p>i sometime have discussions with friends on their phd research work. recently, they give me a few pointers on network design: </p>\n\n<ol>\n<li><p>it is possible that a single batch norm cannot normalized the feature well for multi-task. different task needs different normalization.  Care needs to be taken at designing the head</p></li>\n<li><p>you can force the network to learn high frequency features (e.g.  edge, etc) over texture  by adding loss related to high frequency feature extraction , together with the usual cross entropy classification loss.</p></li>\n</ol>",
      "rawMarkdown": "i sometime have discussions with friends on their phd research work. recently, they give me a few pointers on network design: \n\n1. it is possible that a single batch norm cannot normalized the feature well for multi-task. different task needs different normalization.  Care needs to be taken at designing the head\n\n2. you can force the network to learn high frequency features (e.g.  edge, etc) over texture  by adding loss related to high frequency feature extraction , together with the usual cross entropy classification loss.",
      "votes": 7,
      "replies": [
        {
          "id": 737512,
          "postDate": "2020-02-05T12:57:30.470Z",
          "content": "<p>related:\nusing Adversarial Examples to expand dataset</p>\n\n<p>\"With an enhanced EfficientNet-B8, our method achieves the state-of-the-art 85.5% ImageNet\ntop-1 accuracy without extra data. This result even surpasses the best model in [20] which is trained with 3.5B Instagram images (∼3000× more than ImageNet) and ∼9.4× more parameters.\"</p>\n\n<p>Adversarial Examples Improve Image Recognition\n<a href=\"https://arxiv.org/abs/1911.09665\">https://arxiv.org/abs/1911.09665</a>\n<a href=\"https://www.youtube.com/watch?v=KTCztkNJm50\">https://www.youtube.com/watch?v=KTCztkNJm50</a></p>\n\n<p><img src=\"https://i.redd.it/8rm53y9puf141.png\" alt=\"\"> <img src=\"https://neurohive.io/wp-content/uploads/2019/11/Screenshot-from-2019-11-26-23-34-49-570x419.png\" alt=\"\"></p>\n\n<p>see also:\nCompounding the Performance Improvements of Assembled Techniques in a Convolutional Neural Network <a href=\"https://arxiv.org/pdf/2001.06268.pdf\">https://arxiv.org/pdf/2001.06268.pdf</a></p>",
          "rawMarkdown": "related:\nusing Adversarial Examples to expand dataset\n\n\"With an enhanced EfficientNet-B8, our method achieves the state-of-the-art 85.5% ImageNet\ntop-1 accuracy without extra data. This result even surpasses the best model in [20] which is trained with 3.5B Instagram images (∼3000× more than ImageNet) and ∼9.4× more parameters.\"\n\nAdversarial Examples Improve Image Recognition\nhttps://arxiv.org/abs/1911.09665\nhttps://www.youtube.com/watch?v=KTCztkNJm50\n\n![](https://i.redd.it/8rm53y9puf141.png) ![](https://neurohive.io/wp-content/uploads/2019/11/Screenshot-from-2019-11-26-23-34-49-570x419.png)\n\nsee also:\nCompounding the Performance Improvements of Assembled Techniques in a Convolutional Neural Network https://arxiv.org/pdf/2001.06268.pdf\n\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 723886,
      "postDate": "2020-01-20T15:16:12.777Z",
      "content": "<p>magic net</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd43e2775c784672d47d6b37cc582cb90%2FSelection_073.png?generation=1579533369553774&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd5ff591bb2d20c7c87753e6c2170189b%2FSelection_079.png?generation=1579594728840307&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "magic net\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd43e2775c784672d47d6b37cc582cb90%2FSelection_073.png?generation=1579533369553774&amp;alt=media)\n\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd5ff591bb2d20c7c87753e6c2170189b%2FSelection_079.png?generation=1579594728840307&amp;alt=media)\n",
      "votes": 8,
      "replies": [
        {
          "id": 724702,
          "postDate": "2020-01-21T11:56:12.537Z",
          "content": "<p>single model, single fold performance:\n```\nLB = 0.9755\nCV = 0.983629\n               grapheme_root  0.977137\n             vowel_diacritic  0.993395\n         consonant_diacritic  0.986848 </p>\n\n<p>model: SE-Resnext50 with stride replaced with max pooling + SWA\ninput: 137x236\naugmentation: mixup + cut-fade</p>\n\n<p>```</p>\n\n<p>after more tunning:</p>\n\n<p>```\nLB = 0.9760\nCV = 0.984155\n               grapheme_root  0.978920\n             vowel_diacritic  0.992851\n         consonant_diacritic  0.985929</p>\n\n<p>```</p>",
          "rawMarkdown": "single model, single fold performance:\n```\nLB = 0.9755\nCV = 0.983629\n               grapheme_root  0.977137\n             vowel_diacritic  0.993395\n         consonant_diacritic  0.986848 \n\nmodel: SE-Resnext50 with stride replaced with max pooling + SWA\ninput: 137x236\naugmentation: mixup + cut-fade\n\n```\n\nafter more tunning:\n\n```\nLB = 0.9760\nCV = 0.984155\n               grapheme_root  0.978920\n             vowel_diacritic  0.992851\n         consonant_diacritic  0.985929\n                    \n\n\n```"
        },
        {
          "id": 726108,
          "postDate": "2020-01-22T20:27:47.137Z",
          "content": "<p>Hi Heng! I am trying to automate the 'magic' net architecture. First I replaced de stride 2 to 1, saving the conv 'path':</p>\n\n<p><code>\nmodules = {}\nfor name, module in model.named_modules():\n    if(isinstance(module, nn.Conv2d)):\n        stride = module.stride\n        if stride == (2, 2) or stride == 2:\n            module.stride = (1,1)\n            modules[name] = module\n        elif stride == 2:\n            module.stride = 1\n            modules[name] = module\n</code></p>\n\n<p>Next I try to insert the maxpool as follows:</p>\n\n<p>```\nfor name in modules:\n    parent_module = model\n    objs = name.split(\".\")\n    if len(objs) == 1:\n        #model.<strong>setattr</strong>(name, modules[name])\n        model.<strong>setattr</strong>(\"magicMaxPool\", nn.MaxPool2d(kernel_size=2, stride=2))\n        continue</p>\n\n<pre><code>for obj in objs[:-1]:\n    parent_module = parent_module.__getattr__(obj)\n\nparent_module.__setattr__(\"magicMaxPool\", nn.MaxPool2d(kernel_size=2, stride=2))\n</code></pre>\n\n<p>```</p>\n\n<p>If I inspect the net/modules, it appears the maxpool but when I do the inference stride works properly but maxpool is not used (vector size before global avg pool is too big).</p>",
          "rawMarkdown": "Hi Heng! I am trying to automate the 'magic' net architecture. First I replaced de stride 2 to 1, saving the conv 'path':\n\n```\nmodules = {}\nfor name, module in model.named_modules():\n    if(isinstance(module, nn.Conv2d)):\n        stride = module.stride\n        if stride == (2, 2) or stride == 2:\n            module.stride = (1,1)\n            modules[name] = module\n        elif stride == 2:\n            module.stride = 1\n            modules[name] = module\n```\n\nNext I try to insert the maxpool as follows:\n\n```\nfor name in modules:\n    parent_module = model\n    objs = name.split(\".\")\n    if len(objs) == 1:\n        #model.__setattr__(name, modules[name])\n        model.__setattr__(\"magicMaxPool\", nn.MaxPool2d(kernel_size=2, stride=2))\n        continue\n\n    for obj in objs[:-1]:\n        parent_module = parent_module.__getattr__(obj)\n\n    parent_module.__setattr__(\"magicMaxPool\", nn.MaxPool2d(kernel_size=2, stride=2))\n```\n\nIf I inspect the net/modules, it appears the maxpool but when I do the inference stride works properly but maxpool is not used (vector size before global avg pool is too big).",
          "votes": 1
        },
        {
          "id": 729048,
          "postDate": "2020-01-25T16:46:53.873Z",
          "content": "<p>related: <a href=\"http://openaccess.thecvf.com/content_cvpr_2017/papers/Zhai_S3Pool_Pooling_With_CVPR_2017_paper.pdf\">http://openaccess.thecvf.com/content_cvpr_2017/papers/Zhai_S3Pool_Pooling_With_CVPR_2017_paper.pdf</a></p>\n\n<p>S3Pool: Pooling with Stochastic Spatial Sampling</p>",
          "rawMarkdown": "related: http://openaccess.thecvf.com/content_cvpr_2017/papers/Zhai_S3Pool_Pooling_With_CVPR_2017_paper.pdf\n\nS3Pool: Pooling with Stochastic Spatial Sampling",
          "votes": 2
        },
        {
          "id": 735731,
          "postDate": "2020-02-03T11:53:18.557Z",
          "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> \nI always look at the code you write and I am very grateful.\nI try to training Resnext50 with magic net,But cv value did not go above 0.96. Did you do anything other than change the layer?</p>",
          "rawMarkdown": "Hi @hengck23 \nI always look at the code you write and I am very grateful.\nI try to training Resnext50 with magic net,But cv value did not go above 0.96. Did you do anything other than change the layer?"
        }
      ]
    },
    {
      "id": 707588,
      "postDate": "2020-01-01T05:35:10.807Z",
      "content": "<p>after a series of experiments, how are my conclusion:</p>\n\n<ol>\n<li>final first LB ranking: 0.99+ after 3 months</li>\n<li>shakeup: +/- 0.002 (should be quite small)</li>\n</ol>\n\n<p>baseline results:\n a. 0.9675~0.9715\n  - standard method (joint 3 class) and ensemble standard model (e.g. densenet, resnet, efficientnet, etc) \n  - to maximize performance, search for best input size (e.g. original size, enlarged size like 224x224,256x256, or even larger), try better augmentation, train longer with better learning rate (swa, cyclic annealing, ...), etc\n  - size, scale normalization</p>\n\n<p>b. 0.9700~0.9800\n... to be updated ... (better formulation, loss etc?)</p>\n\n<p>c. 0.9800~0.9900\n... to be updated ... (some way to create more data?)</p>\n\n<p>don't forget  Google Cloud AutoML Vision as part of solution !!!!</p>",
      "rawMarkdown": "after a series of experiments, how are my conclusion:\n\n1. final first LB ranking: 0.99+ after 3 months\n2. shakeup: +/- 0.002 (should be quite small)\n\nbaseline results:\n a. 0.9675~0.9715\n  - standard method (joint 3 class) and ensemble standard model (e.g. densenet, resnet, efficientnet, etc) \n  - to maximize performance, search for best input size (e.g. original size, enlarged size like 224x224,256x256, or even larger), try better augmentation, train longer with better learning rate (swa, cyclic annealing, ...), etc\n  - size, scale normalization\n \n b. 0.9700~0.9800\n... to be updated ... (better formulation, loss etc?)\n\n c. 0.9800~0.9900\n... to be updated ... (some way to create more data?)\n\ndon't forget  Google Cloud AutoML Vision as part of solution !!!!",
      "votes": 8,
      "replies": [
        {
          "id": 708848,
          "postDate": "2020-01-02T19:00:30.557Z",
          "content": "<p>Thanks for the pointers <a href=\"/hengck23\">@hengck23</a>  . I started playing with augmentations . I have got the cutout working here . </p>\n\n<p><a href=\"https://www.kaggle.com/phoenix9032/pytorch-efficientnet-starter-code\">https://www.kaggle.com/phoenix9032/pytorch-efficientnet-starter-code</a></p>\n\n<p>Now , need to see how MixUp and RICAP and other thing can work out along with the training for multi-heads . Need to make some modification from the original implementation.</p>\n\n<p>We can try dual cutout as well. </p>",
          "rawMarkdown": "Thanks for the pointers @hengck23  . I started playing with augmentations . I have got the cutout working here . \n\nhttps://www.kaggle.com/phoenix9032/pytorch-efficientnet-starter-code\n\nNow , need to see how MixUp and RICAP and other thing can work out along with the training for multi-heads . Need to make some modification from the original implementation.\n\nWe can try dual cutout as well. ",
          "votes": 3
        },
        {
          "id": 719435,
          "postDate": "2020-01-15T14:01:09.920Z",
          "content": "<p>some important (?) experimental results</p>\n\n<ul>\n<li><p>sampling affects results (e.g. balance sampling of \"root\" or \"grapheme\" affects \"root\", \"constant\", \"vowel\" differently)</p></li>\n<li><p>augmentation affects results (e.g. some drop in accuracy when scaling and rotation is used)</p></li>\n<li><p>arcface and large-margin like softamx does improve results but cannot be used with mixup?</p></li>\n</ul>",
          "rawMarkdown": "some important (?) experimental results\n\n- sampling affects results (e.g. balance sampling of \"root\" or \"grapheme\" affects \"root\", \"constant\", \"vowel\" differently)\n\n- augmentation affects results (e.g. some drop in accuracy when scaling and rotation is used)\n\n- arcface and large-margin like softamx does improve results but cannot be used with mixup?",
          "votes": 2
        },
        {
          "id": 721090,
          "postDate": "2020-01-17T03:20:43.097Z",
          "content": "<p>Hi Heng <a href=\"/hengck23\">@hengck23</a> (and other Kagglers), I see in your latest starting kit that you added a fourth target of the 1295 unique classes of Bengali letters as provided in the training dataset. If the testing dataset (which is not accessible to us) has letters other than those 1295 classes, would the training approach be prone to overfitting? Please correct me if I am wrong. </p>",
          "rawMarkdown": "Hi Heng @hengck23 (and other Kagglers), I see in your latest starting kit that you added a fourth target of the 1295 unique classes of Bengali letters as provided in the training dataset. If the testing dataset (which is not accessible to us) has letters other than those 1295 classes, would the training approach be prone to overfitting? Please correct me if I am wrong. "
        },
        {
          "id": 721132,
          "postDate": "2020-01-17T04:53:19.677Z",
          "content": "<p>you can use low weight like 0.1 for the fourth class, you can also use 1295+1 (background class for none of the above). to create background class, selectively flip the original train samples</p>",
          "rawMarkdown": "you can use low weight like 0.1 for the fourth class, you can also use 1295+1 (background class for none of the above). to create background class, selectively flip the original train samples"
        }
      ]
    },
    {
      "id": 758757,
      "postDate": "2020-02-28T05:38:01.697Z",
      "content": "<p>good for augmentation?</p>\n\n<p><img src=\"https://raw.githubusercontent.com/warbean/tps_stn_pytorch/master/demo/top_1.gif\" alt=\"\"></p>\n\n<p><a href=\"https://github.com/WarBean/tps_stn_pytorch\">https://github.com/WarBean/tps_stn_pytorch</a></p>",
      "rawMarkdown": "good for augmentation?\n\n![](https://raw.githubusercontent.com/warbean/tps_stn_pytorch/master/demo/top_1.gif)\n\nhttps://github.com/WarBean/tps_stn_pytorch",
      "votes": 5,
      "replies": [
        {
          "id": 759014,
          "postDate": "2020-02-28T13:02:09.523Z",
          "content": "<p>very interesting augmemtation!</p>",
          "rawMarkdown": "very interesting augmemtation!"
        }
      ]
    },
    {
      "id": 724848,
      "postDate": "2020-01-21T14:51:04.107Z",
      "content": "<p>in google doddle, you can add and subtract drawing. I wonder if this can be done in bengali grapheme? e.g. add/remove or replace vowel, etc\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F249a956c438d0dd9ef6f618882dde10a%2FSelection_087.png?generation=1579618261743829&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "in google doddle, you can add and subtract drawing. I wonder if this can be done in bengali grapheme? e.g. add/remove or replace vowel, etc\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F249a956c438d0dd9ef6f618882dde10a%2FSelection_087.png?generation=1579618261743829&amp;alt=media)\n\n\n",
      "votes": 6
    },
    {
      "id": 723868,
      "postDate": "2020-01-20T14:54:31.900Z",
      "content": "<p>new augmentation</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb35fa0d1187ed81acb31b38fedbc6828%2FSelection_069.png?generation=1579532069051469&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "new augmentation\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb35fa0d1187ed81acb31b38fedbc6828%2FSelection_069.png?generation=1579532069051469&amp;alt=media)\n",
      "votes": 6,
      "replies": [
        {
          "id": 724274,
          "postDate": "2020-01-21T01:59:26.787Z",
          "content": "<p>related: augmentation on feature map:</p>\n\n<ol>\n<li>choose some random feature map at training</li>\n<li>choose same max value</li>\n<li>modify: new value = alpha * old value, where alpha is between 0 to 1 (preferably near 0.5)</li>\n</ol>\n\n<p>this attenuate feature values, preventing it from over dominating and lead to over fitting </p>",
          "rawMarkdown": "related: augmentation on feature map:\n\n1. choose some random feature map at training\n2. choose same max value\n3. modify: new value = alpha * old value, where alpha is between 0 to 1 (preferably near 0.5)\n\nthis attenuate feature values, preventing it from over dominating and lead to over fitting "
        }
      ]
    },
    {
      "id": 775716,
      "postDate": "2020-03-17T00:23:49.243Z",
      "content": "<p>private lb of coded posted here:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fda66de721f1d8955b9925e3642f5a0f0%2FSelection_074.png?generation=1584404626455936&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "private lb of coded posted here:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fda66de721f1d8955b9925e3642f5a0f0%2FSelection_074.png?generation=1584404626455936&amp;alt=media)\n",
      "votes": 3,
      "replies": [
        {
          "id": 775910,
          "postDate": "2020-03-17T02:45:02.400Z",
          "content": "<p>Amazing results but sorry to see they are not selected...This post is really helpful and I have learned a lot from your code for the first competition I decided to take part in seriously. Thank you!\nBy the way, did you forget to select final submission?😂  Or just because of the shake up?</p>",
          "rawMarkdown": "Amazing results but sorry to see they are not selected...This post is really helpful and I have learned a lot from your code for the first competition I decided to take part in seriously. Thank you!\nBy the way, did you forget to select final submission?😂  Or just because of the shake up?"
        }
      ]
    },
    {
      "id": 765573,
      "postDate": "2020-03-06T20:08:44.370Z",
      "content": "<p>there is one extra supervision signal.</p>\n\n<p><code>\n list('ক্ট্রো')\n['ক', '্', 'ট', '্', 'র', 'ো']\n</code>\na list command breaks the graheme into more component. you can use this the measure the amount of common parts (distance) between 2 graphemes </p>",
      "rawMarkdown": "there is one extra supervision signal.\n\n```\n list('ক্ট্রো')\n['ক', '্', 'ট', '্', 'র', 'ো']\n```\na list command breaks the graheme into more component. you can use this the measure the amount of common parts (distance) between 2 graphemes ",
      "votes": 3
    },
    {
      "id": 761546,
      "postDate": "2020-03-02T16:38:15.470Z",
      "content": "<p>if this model is not correlated to image model, it is good for ensemble.\nthis work enables input of very large image</p>\n\n<p>Learning in the Frequency Domain\nKai Xu, Minghai Qin, Fei Sun, Yuhao Wang, Yen-kuang Chen, Fengbo Ren</p>\n\n<p><img src=\"https://storage.googleapis.com/groundai-web-prod/media/users/user_12457/project_409932/images/x2.png\" alt=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/groundai-web-prod/media/users/user_12457/project_409932/images/x4.png\" alt=\"\"></p>\n\n<p>```</p>\n\n<p>ResNet-50   #Channels   Size Per Channel    Top-1   Top-5   Normalized Input Size <br>\nRGB             3               224x224                   75.780    92.650  1.0 <br>\nDCT-24 (ours)    24             56x56                     77.196    93.504  0.5      </p>\n\n<p>```</p>",
      "rawMarkdown": "if this model is not correlated to image model, it is good for ensemble.\nthis work enables input of very large image\n\nLearning in the Frequency Domain\nKai Xu, Minghai Qin, Fei Sun, Yuhao Wang, Yen-kuang Chen, Fengbo Ren\n\n![](https://storage.googleapis.com/groundai-web-prod/media/users/user_12457/project_409932/images/x2.png)\n\n![](https://storage.googleapis.com/groundai-web-prod/media/users/user_12457/project_409932/images/x4.png)\n\n\n```\n\nResNet-50\t#Channels\tSize Per Channel\tTop-1\tTop-5\tNormalized Input Size\t\nRGB\t            3\t            224x224\t                  75.780\t92.650\t1.0\t \t \nDCT-24 (ours)    24\t            56x56\t                  77.196\t93.504\t0.5\t \t \n \n\n\n```",
      "votes": 3,
      "replies": [
        {
          "id": 763297,
          "postDate": "2020-03-04T10:37:02.583Z",
          "content": "<p>Thanks for sharing an interesting paper.\nThe code is available below.</p>\n\n<p><a href=\"https://github.com/calmevtime1990/supp/tree/master/classification\">https://github.com/calmevtime1990/supp/tree/master/classification</a></p>",
          "rawMarkdown": "Thanks for sharing an interesting paper.\nThe code is available below.\n\nhttps://github.com/calmevtime1990/supp/tree/master/classification",
          "votes": 1
        }
      ]
    },
    {
      "id": 728658,
      "postDate": "2020-01-25T03:28:10.440Z",
      "content": "<p>top-1 and top-2 recall:</p>\n\n<p>```\nlocal validation results</p>\n\n<p>avgerage recall (top-1): 0.982253\n               grapheme_root  0.974616\n             vowel_diacritic  0.992676\n         consonant_diacritic  0.987103</p>\n\n<p>avgerage recall (top-2) : 0.995918\n               grapheme_root  0.993451\n             vowel_diacritic  0.998653\n         consonant_diacritic  0.998116</p>\n\n<p>```</p>\n\n<p>if you make a mistake, the correct results is probably the top-2. if you can think of a good way to post-process, you can improve your score</p>",
      "rawMarkdown": "top-1 and top-2 recall:\n\n```\nlocal validation results\n\navgerage recall (top-1): 0.982253\n               grapheme_root  0.974616\n             vowel_diacritic  0.992676\n         consonant_diacritic  0.987103\n                    \n\navgerage recall (top-2) : 0.995918\n               grapheme_root  0.993451\n             vowel_diacritic  0.998653\n         consonant_diacritic  0.998116\n                   \n\n```\n\nif you make a mistake, the correct results is probably the top-2. if you can think of a good way to post-process, you can improve your score",
      "votes": 3,
      "replies": [
        {
          "id": 729149,
          "postDate": "2020-01-25T20:10:12.520Z",
          "content": "<p>this screams for an ensemble of models ;)</p>",
          "rawMarkdown": "this screams for an ensemble of models ;)"
        }
      ]
    },
    {
      "id": 757812,
      "postDate": "2020-02-27T05:58:26.333Z",
      "content": "<p>New paper today\nOn Feature Normalization and Data Augmentation</p>\n\n<p><a href=\"https://arxiv.org/pdf/2002.11102.pdf\">https://arxiv.org/pdf/2002.11102.pdf</a></p>",
      "rawMarkdown": "New paper today\nOn Feature Normalization and Data Augmentation\n\nhttps://arxiv.org/pdf/2002.11102.pdf",
      "votes": 4,
      "replies": [
        {
          "id": 757817,
          "postDate": "2020-02-27T06:09:36.200Z",
          "content": "<p>arXiv:2002.11022 (cross-list from cs.LG) [pdf, other]\nBeyond Dropout: Feature Map Distortion to Regularize Deep Neural Networks\nYehui Tang, Yunhe Wang, Yixing Xu, Boxin Shi, Chao Xu, Chunjing Xu, Chang Xu\nSubjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)</p>",
          "rawMarkdown": "arXiv:2002.11022 (cross-list from cs.LG) [pdf, other]\nBeyond Dropout: Feature Map Distortion to Regularize Deep Neural Networks\nYehui Tang, Yunhe Wang, Yixing Xu, Boxin Shi, Chao Xu, Chunjing Xu, Chang Xu\nSubjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)",
          "votes": 1
        }
      ]
    },
    {
      "id": 706574,
      "postDate": "2019-12-30T15:27:50.050Z",
      "content": "<p>What's the meaning of “without bn refinement”? </p>",
      "rawMarkdown": "What's the meaning of “without bn refinement”? ",
      "votes": 3,
      "replies": [
        {
          "id": 706611,
          "postDate": "2019-12-30T16:07:09.213Z",
          "content": "<p>I think Heng is talking about the recently proposed 'Full Normalization' technique, a better alternative (or you can say modification) to the conventional Batch Normalization. The idea is described in the following paper:\n<a href=\"https://arxiv.org/abs/1810.06177\">https://arxiv.org/abs/1810.06177</a></p>",
          "rawMarkdown": "I think Heng is talking about the recently proposed 'Full Normalization' technique, a better alternative (or you can say modification) to the conventional Batch Normalization. The idea is described in the following paper:\nhttps://arxiv.org/abs/1810.06177"
        },
        {
          "id": 706618,
          "postDate": "2019-12-30T16:11:02.760Z",
          "content": "<p><a href=\"https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/\">https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/</a></p>\n\n<p>```\nBATCH NORMALIZATION\nOne important detail to keep in mind is batch normalization. Batch normalization layers compute running statistics of activations during training. Note that the SWA averages of the weights are never used to make predictions during training, and so the batch normalization layers do not have the activation statistics computed after you reset the weights of your model with opt.swap_swa_sgd(). To compute the activation statistics you can just make a forward pass on your training data using the SWA model once the training is finished. In the SWA class we provide a helper function opt.bn_update(train_loader, model). It updates the activation statistics for every batch normalization layer in the model by making a forward pass on the train_loader data loader. You only need to call this function once in the end of training.</p>\n\n<p>```</p>\n\n<p>the above step is not implemented</p>",
          "rawMarkdown": "https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/\n\n```\nBATCH NORMALIZATION\nOne important detail to keep in mind is batch normalization. Batch normalization layers compute running statistics of activations during training. Note that the SWA averages of the weights are never used to make predictions during training, and so the batch normalization layers do not have the activation statistics computed after you reset the weights of your model with opt.swap_swa_sgd(). To compute the activation statistics you can just make a forward pass on your training data using the SWA model once the training is finished. In the SWA class we provide a helper function opt.bn_update(train_loader, model). It updates the activation statistics for every batch normalization layer in the model by making a forward pass on the train_loader data loader. You only need to call this function once in the end of training.\n\n```\n\nthe above step is not implemented",
          "votes": 5
        },
        {
          "id": 706626,
          "postDate": "2019-12-30T16:19:48.630Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  Thanks for your clarification! </p>",
          "rawMarkdown": "@hengck23  Thanks for your clarification! "
        }
      ]
    },
    {
      "id": 706492,
      "postDate": "2019-12-30T13:15:02.323Z",
      "content": "<p>What's the meaning of swa ?</p>",
      "rawMarkdown": "What's the meaning of swa ?",
      "votes": 3,
      "replies": [
        {
          "id": 706607,
          "postDate": "2019-12-30T16:03:32.357Z",
          "content": "<p>Stochastic Weight Averaging, you can read this paper for more details:\n<a href=\"https://arxiv.org/abs/1803.05407\">https://arxiv.org/abs/1803.05407</a></p>",
          "rawMarkdown": "Stochastic Weight Averaging, you can read this paper for more details:\nhttps://arxiv.org/abs/1803.05407",
          "votes": 4
        },
        {
          "id": 707056,
          "postDate": "2019-12-31T08:05:01.543Z",
          "content": "<p>thx</p>",
          "rawMarkdown": "thx",
          "votes": 1
        }
      ]
    },
    {
      "id": 728668,
      "postDate": "2020-01-25T04:01:19.887Z",
      "content": "<p>multi-task and task dependency</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fefb21bf6c5f708e84efb92a2a610f24f%2FClipboard05.png?generation=1579924874232793&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0b65fbdbc95cf5730ca99eff92e7bf6f%2FClipboard04.png?generation=1579924877627631&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "multi-task and task dependency\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fefb21bf6c5f708e84efb92a2a610f24f%2FClipboard05.png?generation=1579924874232793&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0b65fbdbc95cf5730ca99eff92e7bf6f%2FClipboard04.png?generation=1579924877627631&amp;alt=media)\n",
      "votes": 4,
      "replies": [
        {
          "id": 740985,
          "postDate": "2020-02-10T03:46:02.093Z",
          "content": "<p>Im studying the correlation among tasks. What's this paper? Thanks!</p>",
          "rawMarkdown": "Im studying the correlation among tasks. What's this paper? Thanks!"
        }
      ]
    },
    {
      "id": 726161,
      "postDate": "2020-01-22T21:58:56.863Z",
      "content": "<p>loss formulation:</p>\n\n<p>some task (root, constant, or vowel) are more confidence than others.\nsimilarly, some classes are easier within a class.</p>\n\n<p>e.g. given task 1 has class A,B,C,D,E and task 2 has class a,b,c and if we know that some class cannot co-exist, it may give better results if we apply the rules as follow:\n```\n1. naive method:\ntreat task 1 as 5 classes  problem and task 2 as 3 classes. ignore any co-occurrence prior knowledge</p>\n\n<ol>\n<li>use prior knowledge\nif input is a, then it can only be A,B,C\nif input is b, then it can only be A,B,E\nif input is c, then it can only be D,E</li>\n</ol>\n\n<p>```\nif you know the language, you may apply such rules</p>",
      "rawMarkdown": "loss formulation:\n\nsome task (root, constant, or vowel) are more confidence than others.\nsimilarly, some classes are easier within a class.\n\ne.g. given task 1 has class A,B,C,D,E and task 2 has class a,b,c and if we know that some class cannot co-exist, it may give better results if we apply the rules as follow:\n```\n1. naive method:\ntreat task 1 as 5 classes  problem and task 2 as 3 classes. ignore any co-occurrence prior knowledge\n\n 2. use prior knowledge\nif input is a, then it can only be A,B,C\nif input is b, then it can only be A,B,E\nif input is c, then it can only be D,E\n\n```\nif you know the language, you may apply such rules\n",
      "votes": 4,
      "replies": [
        {
          "id": 726925,
          "postDate": "2020-01-23T10:16:45.180Z",
          "content": "<p>or we could change the architecture of model's head such that at the time of predicting task 2 we feed the probabilities(softmax) for task 1 predicted by our model.</p>",
          "rawMarkdown": "or we could change the architecture of model's head such that at the time of predicting task 2 we feed the probabilities(softmax) for task 1 predicted by our model.",
          "votes": 2
        }
      ]
    },
    {
      "id": 716656,
      "postDate": "2020-01-12T04:17:58.407Z",
      "content": "<p>version 20200111:\nse-resnext50 + balanced sampler + mixup (please refer to readme.ppt at the google drive)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F97673c14b01d4aff29fc1d688b036267%2FSelection_069.png?generation=1578830084131882&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "version 20200111:\nse-resnext50 + balanced sampler + mixup (please refer to readme.ppt at the google drive)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F97673c14b01d4aff29fc1d688b036267%2FSelection_069.png?generation=1578830084131882&amp;alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 717189,
          "postDate": "2020-01-12T21:04:28.410Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> thanks for your sharing again. Do you have the updated verison of <code>compute_kaggle_metric</code> in your <code>kaggle.py</code>?</p>",
          "rawMarkdown": "@hengck23 thanks for your sharing again. Do you have the updated verison of `compute_kaggle_metric` in your `kaggle.py`?"
        },
        {
          "id": 717592,
          "postDate": "2020-01-13T10:50:10.053Z",
          "content": "<p>```\ndef compute_kaggle_metric(probability, truth):</p>\n\n<pre><code>def compute_recall(probability,truth):\n    num_class = probability.shape[-1]\n    y = probability.argmax(-1)\n    t = truth\n    correct = y==t\n\n    recall = np.zeros(num_class)\n    for c in range(num_class):\n        e = correct[t==c]\n        if len(e)&amp;gt;0:\n            recall[c]=e.mean()\n    return recall\n\ncomponet = []\nrecall   = []\nfor p,t in zip(probability,truth):\n    r = compute_recall(p,t)\n    recall.append(r)\n    componet.append(r.mean())\n\naverage = np.average(componet, weights=[2,1,1,0])\nreturn average, componet, recall\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "```\ndef compute_kaggle_metric(probability, truth):\n\n    def compute_recall(probability,truth):\n        num_class = probability.shape[-1]\n        y = probability.argmax(-1)\n        t = truth\n        correct = y==t\n\n        recall = np.zeros(num_class)\n        for c in range(num_class):\n            e = correct[t==c]\n            if len(e)&gt;0:\n                recall[c]=e.mean()\n        return recall\n\n    componet = []\n    recall   = []\n    for p,t in zip(probability,truth):\n        r = compute_recall(p,t)\n        recall.append(r)\n        componet.append(r.mean())\n\n    average = np.average(componet, weights=[2,1,1,0])\n    return average, componet, recall\n```",
          "votes": 2
        }
      ]
    },
    {
      "id": 706712,
      "postDate": "2019-12-30T18:06:51.830Z",
      "content": "<p>there are 168 grapheme_root, 11 vowel_diacritic, 7 consonant_diacritic and 1295 grapheme classes!\nyou can build a fourth classifier for 1295 grapheme classes and decode the component from it</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbeafecfbcd56e0f8c5ee4ed6d98bd2b3%2FSelection_088.png?generation=1577729205397481&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0e6843ab6fef11c4b7b478c45d0f07f2%2FSelection_087.png?generation=1577729209517703&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "there are 168 grapheme\\_root, 11 vowel\\_diacritic, 7 consonant\\_diacritic and 1295 grapheme classes!\nyou can build a fourth classifier for 1295 grapheme classes and decode the component from it\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbeafecfbcd56e0f8c5ee4ed6d98bd2b3%2FSelection_088.png?generation=1577729205397481&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0e6843ab6fef11c4b7b478c45d0f07f2%2FSelection_087.png?generation=1577729209517703&amp;alt=media)\n",
      "votes": 4,
      "replies": [
        {
          "id": 706932,
          "postDate": "2019-12-31T03:04:58.270Z",
          "content": "<p>Find a way to balance per class accuracy. &lt;- I thought this approach too and tried a bit although I don’t have convincing reason why it is helpful. Do you have it?</p>",
          "rawMarkdown": "Find a way to balance per class accuracy. &lt;- I thought this approach too and tried a bit although I don’t have convincing reason why it is helpful. Do you have it?"
        },
        {
          "id": 717364,
          "postDate": "2020-01-13T03:58:19.630Z",
          "rawMarkdown": ""
        },
        {
          "id": 759126,
          "postDate": "2020-02-28T16:12:54.400Z",
          "content": "<p>I was curious about how to build a fourth classifier for 1295 grapheme classes,is that just groupby grapheme and label it from 1 to 1295?\nHope you could clarify,thanks.</p>",
          "rawMarkdown": "I was curious about how to build a fourth classifier for 1295 grapheme classes,is that just groupby grapheme and label it from 1 to 1295?\nHope you could clarify,thanks."
        },
        {
          "id": 759134,
          "postDate": "2020-02-28T16:19:22.130Z",
          "content": "<p><a href=\"/thefatcat\">@thefatcat</a> </p>\n\n<p>yes, you are correct</p>",
          "rawMarkdown": "@thefatcat \n\nyes, you are correct"
        }
      ]
    },
    {
      "id": 706395,
      "postDate": "2019-12-30T10:31:29.073Z",
      "content": "<p>augmentation in the release 'version 20191230'</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F13e81e022c9b8ebf067517e15d62e938%2FSelection_081.png?generation=1577701870877750&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "augmentation in the release 'version 20191230'\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F13e81e022c9b8ebf067517e15d62e938%2FSelection_081.png?generation=1577701870877750&amp;alt=media)\n",
      "votes": 4,
      "replies": [
        {
          "id": 706738,
          "postDate": "2019-12-30T18:55:33.683Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  Hi. Can u tell what augmentations u used ?  Thanks</p>",
          "rawMarkdown": "@hengck23  Hi. Can u tell what augmentations u used ?  Thanks"
        }
      ]
    },
    {
      "id": 716020,
      "postDate": "2020-01-11T06:47:03.763Z",
      "content": "<p>I published a <a href=\"https://www.kaggle.com/bibek777/heng-s-starter-training-kernel\">kernel</a> based on this so that people can train it using kaggle GPUs</p>",
      "rawMarkdown": "I published a [kernel](https://www.kaggle.com/bibek777/heng-s-starter-training-kernel) based on this so that people can train it using kaggle GPUs",
      "votes": 2
    },
    {
      "id": 760615,
      "postDate": "2020-03-01T14:25:04.707Z",
      "content": "<p>yet another paper:\nMaxUp: A Simple Way to Improve Generalization of Neural Network Training\n```</p>\n\n<p>Method  Top-1 error Top-5 error\nVanilla (He et al., 2016a)  76.3    -\nDropout (Srivastava et al., 2014)   76.8    93.4\nDropPath (Larsson et al., 2017) 77.1    93.5\nManifold Mixup (Verma et al., 2019) 77.5    93.8\nAutoAugment (Cubuk et al., 2019a)   77.6    93.8\nMixup (Zhang et al., 2018)  77.9    93.9\nDropBlock (Ghiasi et al., 2018) 78.3    94.1\nCutMix (Yun et al., 2019)   78.6    94.0\nMaxUp+CutMix    78.9    94.2</p>\n\n<p>Table 1: Summary of top1 and top5 accuracies on the validation set of ImageNet for ResNet-50.\n```</p>",
      "rawMarkdown": "yet another paper:\nMaxUp: A Simple Way to Improve Generalization of Neural Network Training\n```\n\nMethod\tTop-1 error\tTop-5 error\nVanilla (He et al., 2016a)\t76.3\t-\nDropout (Srivastava et al., 2014)\t76.8\t93.4\nDropPath (Larsson et al., 2017)\t77.1\t93.5\nManifold Mixup (Verma et al., 2019)\t77.5\t93.8\nAutoAugment (Cubuk et al., 2019a)\t77.6\t93.8\nMixup (Zhang et al., 2018)\t77.9\t93.9\nDropBlock (Ghiasi et al., 2018)\t78.3\t94.1\nCutMix (Yun et al., 2019)\t78.6\t94.0\nMaxUp+CutMix\t78.9\t94.2\n\nTable 1: Summary of top1 and top5 accuracies on the validation set of ImageNet for ResNet-50.\n```",
      "votes": 1
    },
    {
      "id": 758761,
      "postDate": "2020-02-28T05:43:32.973Z",
      "content": "<p>it is possible to treat it as a detection problem\n<img src=\"https://raw.githubusercontent.com/RubanSeven/CRAFT_keras/master/images/CRAFT%E9%AB%98%E6%96%AF%E7%83%AD%E5%9B%BE.png\" alt=\"\"></p>\n\n<p><img src=\"https://pythonawesome.com/content/images/2019/06/CRAFT.jpg\" alt=\"\"></p>",
      "rawMarkdown": "it is possible to treat it as a detection problem\n![](https://raw.githubusercontent.com/RubanSeven/CRAFT_keras/master/images/CRAFT%E9%AB%98%E6%96%AF%E7%83%AD%E5%9B%BE.png)\n\n![](https://pythonawesome.com/content/images/2019/06/CRAFT.jpg)",
      "votes": 1,
      "replies": [
        {
          "id": 759015,
          "postDate": "2020-02-28T13:02:58.843Z",
          "content": "<p>Didn't expect to see CRAFT here</p>",
          "rawMarkdown": "Didn't expect to see CRAFT here"
        },
        {
          "id": 759082,
          "postDate": "2020-02-28T14:49:24.743Z",
          "content": "<p>just an idea\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2e6d456f54d997d4cfabfd1ad5a6af41%2FSelection_082.png?generation=1582901362622740&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "just an idea\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2e6d456f54d997d4cfabfd1ad5a6af41%2FSelection_082.png?generation=1582901362622740&amp;alt=media)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 758090,
      "postDate": "2020-02-27T12:35:38.497Z",
      "content": "<p><a href=\"https://github.com/sujatasaini/Kuzushiji-DropBlock\">https://github.com/sujatasaini/Kuzushiji-DropBlock</a></p>\n\n<p>|Models                           | MNIST | Fashion-MNIST | Kuzushiji-MNIST | Kuzushiji-49 |\n|---------------------------------|-------|---------------|-----------------|--------------|\n|DCNN-DropBlock     | <strong>99.47%</strong> | <strong>93.40%</strong> | <strong>97.66%</strong> | <strong>95.67%</strong> |\n|DCNN-Dropout                         | 97.99% | 85.47% | 86.43% | 95.34% |\n|DCNN-Spatial-Dropout          | 97.17% | 84.44% |  81.08% | 58.18 |</p>",
      "rawMarkdown": "https://github.com/sujatasaini/Kuzushiji-DropBlock\n\n \n\n|Models                           | MNIST | Fashion-MNIST | Kuzushiji-MNIST | Kuzushiji-49 |\n|---------------------------------|-------|---------------|-----------------|--------------|\n|[DCNN-DropBlock](DropBlock/Kuzushiji-MNIST/train.py)     | **99.47%** | **93.40%** | **97.66%** | **95.67%** |\n|[DCNN-Dropout](Dropout/train.py)                         | 97.99% | 85.47% | 86.43% | 95.34% |\n|[DCNN-Spatial-Dropout](SpatialDropout/train.py)          | 97.17% | 84.44% |  81.08% | 58.18 |\n ",
      "votes": 1
    },
    {
      "id": 751629,
      "postDate": "2020-02-20T10:55:19.673Z",
      "content": "<p>another augmentation that works for me:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F5800610b7d3b4f3f03d8e9675b889030%2FSelection_085.png?generation=1582196108850096&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"http://yann.lecun.com/ex/images/invar.png\" alt=\"\"></p>",
      "rawMarkdown": "another augmentation that works for me:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F5800610b7d3b4f3f03d8e9675b889030%2FSelection_085.png?generation=1582196108850096&amp;alt=media)\n\n![](http://yann.lecun.com/ex/images/invar.png)",
      "votes": 1,
      "replies": [
        {
          "id": 751631,
          "postDate": "2020-02-20T10:57:03.813Z",
          "content": "<p>```\ndef train_augment(image, label, infor):\n    original = image.copy()\n    if 1:\n        operation = [\n            lambda image : do_random_erode(image),\n            lambda image : do_random_dilate(image),\n            lambda image : do_random_grid_distortion(image, distort=0.10, num_step = 5),\n            lambda image : do_random_custom_distortion1(image, distort=0.10),\n            lambda image : do_random_crop_rotate_rescale(image, mode={'rotate':10}),\n            lambda image : do_random_block_fade(image, size=[0.1, 0.9], alpha=(0.1,0.5)),\n            lambda image : do_random_line(image),\n        ]\n        num_op = np.random.choice(2)\n        if num_op&gt;0:\n            for op in np.random.choice(operation,num_op):\n                image = op(image)\n        image = do_random_pad_crop(image, 3)</p>\n\n<pre><code>#----------------------------------------------------\nif 1:\n    operation = [\n        lambda image : image,\n        lambda image : do_random_salt_pepper(image),\n        lambda image : do_random_noise(image),\n    ]\n    for op in np.random.choice(operation,1):\n        image = op(image)\n\n#----------------------------------------------------\nreturn original, image, label, infor\n</code></pre>\n\n<p>def train_batch_augment(original, input, onehot):\n    operation = [\n        lambda input, onehot : (input, onehot),\n        lambda input, onehot : do_random_batch_cutout(input, onehot, fill=0),\n        lambda input, onehot : do_random_batch_mixup(original, onehot),\n    ]\n    op = np.random.choice(operation)\n    with torch.no_grad():\n        input, onehot = op(input, onehot)</p>\n\n<pre><code>return input, onehot\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "```\ndef train_augment(image, label, infor):\n    original = image.copy()\n    if 1:\n        operation = [\n            lambda image : do_random_erode(image),\n            lambda image : do_random_dilate(image),\n            lambda image : do_random_grid_distortion(image, distort=0.10, num_step = 5),\n            lambda image : do_random_custom_distortion1(image, distort=0.10),\n            lambda image : do_random_crop_rotate_rescale(image, mode={'rotate':10}),\n            lambda image : do_random_block_fade(image, size=[0.1, 0.9], alpha=(0.1,0.5)),\n            lambda image : do_random_line(image),\n        ]\n        num_op = np.random.choice(2)\n        if num_op&gt;0:\n            for op in np.random.choice(operation,num_op):\n                image = op(image)\n        image = do_random_pad_crop(image, 3)\n\n    #----------------------------------------------------\n    if 1:\n        operation = [\n            lambda image : image,\n            lambda image : do_random_salt_pepper(image),\n            lambda image : do_random_noise(image),\n        ]\n        for op in np.random.choice(operation,1):\n            image = op(image)\n\n    #----------------------------------------------------\n    return original, image, label, infor\n\n\n\ndef train_batch_augment(original, input, onehot):\n    operation = [\n        lambda input, onehot : (input, onehot),\n        lambda input, onehot : do_random_batch_cutout(input, onehot, fill=0),\n        lambda input, onehot : do_random_batch_mixup(original, onehot),\n    ]\n    op = np.random.choice(operation)\n    with torch.no_grad():\n        input, onehot = op(input, onehot)\n\n    return input, onehot\n\n```",
          "votes": 1
        },
        {
          "id": 761929,
          "postDate": "2020-03-03T03:36:04.637Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 737055,
      "postDate": "2020-02-04T21:26:15.080Z",
      "content": "<p>This discussion forum is a gold-mine of ideas. Thanks a lot <a href=\"/hengck23\">@hengck23</a> for initiating this and for sharing such wonderful ideas. </p>\n\n<p>Lots to do, lots of learn, not enough time.</p>",
      "rawMarkdown": "This discussion forum is a gold-mine of ideas. Thanks a lot @hengck23 for initiating this and for sharing such wonderful ideas. \n\nLots to do, lots of learn, not enough time.",
      "votes": 1
    },
    {
      "id": 733502,
      "postDate": "2020-01-31T08:15:44.307Z",
      "content": "<p>don't forget that each of has has addition 36 unlabelled training images at the test parquet files</p>",
      "rawMarkdown": "don't forget that each of has has addition 36 unlabelled training images at the test parquet files",
      "votes": 1
    },
    {
      "id": 733430,
      "postDate": "2020-01-31T05:30:36.087Z",
      "content": "<p>some tricks from ICDAR compeitions</p>\n\n<p><a href=\"https://gateway.newton.ac.uk/sites/default/files/asset/doc/1811/20181127-Large%20Scale%20and%20Unconstrained%20OHCCR%20Based%20on%20Deep%20Learning%20and%20Path%20Signature.pdf\">https://gateway.newton.ac.uk/sites/default/files/asset/doc/1811/20181127-Large%20Scale%20and%20Unconstrained%20OHCCR%20Based%20on%20Deep%20Learning%20and%20Path%20Signature.pdf</a></p>",
      "rawMarkdown": "some tricks from ICDAR compeitions\n\nhttps://gateway.newton.ac.uk/sites/default/files/asset/doc/1811/20181127-Large%20Scale%20and%20Unconstrained%20OHCCR%20Based%20on%20Deep%20Learning%20and%20Path%20Signature.pdf",
      "votes": 1,
      "replies": [
        {
          "id": 733436,
          "postDate": "2020-01-31T05:52:45.607Z",
          "content": "<p>handwriting trajectory recovery\n<img src=\"https://encrypted-tbn0.gstatic.com/images?q=tbn%3AANd9GcQ_0pEXIOBFOypFTak5Pq9drW8LvXBx_W0EdDOoAC4jDTyYO3mf\" alt=\"\"></p>",
          "rawMarkdown": "handwriting trajectory recovery\n![](https://encrypted-tbn0.gstatic.com/images?q=tbn%3AANd9GcQ_0pEXIOBFOypFTak5Pq9drW8LvXBx_W0EdDOoAC4jDTyYO3mf)",
          "votes": 2
        },
        {
          "id": 733846,
          "postDate": "2020-01-31T15:20:33.237Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1001950b2fa0efdac57746685d495781%2FSelection_066.png?generation=1580484022666591&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://arxiv.org/pdf/1606.05763.pdf\">https://arxiv.org/pdf/1606.05763.pdf</a></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1001950b2fa0efdac57746685d495781%2FSelection_066.png?generation=1580484022666591&amp;alt=media)\n\nhttps://arxiv.org/pdf/1606.05763.pdf"
        },
        {
          "id": 733850,
          "postDate": "2020-01-31T15:24:05.603Z",
          "content": "<p>dropsample --- remove outlier\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F631f809daaa2cdbbbe2938d1efd1b8ed%2FSelection_067.png?generation=1580484242967773&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "dropsample --- remove outlier\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F631f809daaa2cdbbbe2938d1efd1b8ed%2FSelection_067.png?generation=1580484242967773&amp;alt=media)\n"
        },
        {
          "id": 733853,
          "postDate": "2020-01-31T15:30:55.200Z",
          "content": "<p>drop-distortion\n<a href=\"https://arxiv.org/pdf/1702.07508.pdf\">https://arxiv.org/pdf/1702.07508.pdf</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fdbe0c609d2f6e881203780e458191c35%2FSelection_068.png?generation=1580484634022137&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "drop-distortion\nhttps://arxiv.org/pdf/1702.07508.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fdbe0c609d2f6e881203780e458191c35%2FSelection_068.png?generation=1580484634022137&amp;alt=media)\n"
        }
      ]
    },
    {
      "id": 724568,
      "postDate": "2020-01-21T08:59:39.887Z",
      "content": "<p>artificial handwriting\n<a href=\"https://distill.pub/2016/handwriting/\">https://distill.pub/2016/handwriting/</a>\n<a href=\"http://otoro.net/ml/\">http://otoro.net/ml/</a>\n<img src=\"http://blog.otoro.net/wp-content/uploads/sites/2/2015/12/cover_simple_full.svg\" alt=\"\"></p>\n\n<p><a href=\"https://genekogan.com/works/a-book-from-the-sky/\">https://genekogan.com/works/a-book-from-the-sky/</a></p>\n\n<p><img src=\"https://genekogan.com/images/a-book-from-the-sky/radinterpolations/rad30is.gif\" alt=\"\">\n<img src=\"https://cdn.rawgit.com/hardmaru/resnet-cppn-gan-tensorflow/master/examples/example_sinusoid.gif\" alt=\"\"></p>",
      "rawMarkdown": "artificial handwriting\nhttps://distill.pub/2016/handwriting/\nhttp://otoro.net/ml/\n![](http://blog.otoro.net/wp-content/uploads/sites/2/2015/12/cover_simple_full.svg)\n\nhttps://genekogan.com/works/a-book-from-the-sky/\n\n ![](https://genekogan.com/images/a-book-from-the-sky/radinterpolations/rad30is.gif)\n![](https://cdn.rawgit.com/hardmaru/resnet-cppn-gan-tensorflow/master/examples/example_sinusoid.gif)",
      "votes": 1
    },
    {
      "id": 724419,
      "postDate": "2020-01-21T06:17:17.547Z",
      "content": "<p><a href=\"https://github.com/osmr/imgclsmob\">https://github.com/osmr/imgclsmob</a>\nmega list of pretrain models</p>",
      "rawMarkdown": "https://github.com/osmr/imgclsmob\nmega list of pretrain models",
      "votes": 1
    },
    {
      "id": 722316,
      "postDate": "2020-01-18T11:44:21.563Z",
      "content": "<p>something interesting ...\n<a href=\"https://github.com/Natsu6767/Generating-Devanagari-Using-DRAW\">https://github.com/Natsu6767/Generating-Devanagari-Using-DRAW</a>\n<img src=\"https://github.com/Natsu6767/Generating-Devanagari-Using-DRAW/raw/master/images/devanagari_generate.gif\" alt=\"\"></p>",
      "rawMarkdown": "something interesting ...\nhttps://github.com/Natsu6767/Generating-Devanagari-Using-DRAW\n![](https://github.com/Natsu6767/Generating-Devanagari-Using-DRAW/raw/master/images/devanagari_generate.gif)",
      "votes": 1,
      "replies": [
        {
          "id": 722322,
          "postDate": "2020-01-18T11:50:32.477Z",
          "content": "<p>Is it like GAN with RNN ? I saw an encoder and decoder structure .This is extremely interesting .</p>",
          "rawMarkdown": "Is it like GAN with RNN ? I saw an encoder and decoder structure .This is extremely interesting .",
          "votes": 1
        },
        {
          "id": 724852,
          "postDate": "2020-01-21T14:55:07.983Z",
          "content": "<p>font RNN\n<a href=\"https://xiazeqing.github.io/FontRNN/\">https://xiazeqing.github.io/FontRNN/</a></p>",
          "rawMarkdown": "font RNN\nhttps://xiazeqing.github.io/FontRNN/"
        }
      ]
    },
    {
      "id": 721139,
      "postDate": "2020-01-17T05:04:47.173Z",
      "content": "<p>one weird idea to fight overfitting:\n```\n1. given  N class train images\n2. flip the images to make 2 N class (if can do transpose, horizontal, vertical flip to give more class)\n3. pretrain a deep network with 2N class</p>\n\n<p>either \n 4. finetune to original N class</p>\n\n<p>or\n4. continue to 2N class and apply test-time augment (TTA) for test + flip images </p>\n\n<p>```</p>\n\n<p>since the problem is now more complex, maybe the network is less likely to overfit</p>",
      "rawMarkdown": "one weird idea to fight overfitting:\n```\n1. given  N class train images\n2. flip the images to make 2 N class (if can do transpose, horizontal, vertical flip to give more class)\n3. pretrain a deep network with 2N class\n\neither \n 4. finetune to original N class\n\nor\n4. continue to 2N class and apply test-time augment (TTA) for test + flip images \n\n```\n\nsince the problem is now more complex, maybe the network is less likely to overfit",
      "votes": 1
    },
    {
      "id": 706961,
      "postDate": "2019-12-31T04:48:04.067Z",
      "content": "<p>In your code you load the dataset based on <code>split=balance2/train_b_fold0_184855.npy</code>. Could you clarify on this? Does it mean that the split fold contains enough samples from each class?</p>",
      "rawMarkdown": "In your code you load the dataset based on `split=balance2/train_b_fold0_184855.npy`. Could you clarify on this? Does it mean that the split fold contains enough samples from each class?",
      "votes": 1,
      "replies": [
        {
          "id": 708134,
          "postDate": "2020-01-02T02:52:50.247Z",
          "content": "<p><a href=\"/bibek777\">@bibek777</a> <a href=\"/hengck23\">@hengck23</a> I am curious how this split is made too. In the previous competition steel and this competition, the split assignment has always been a puzzle. I think it does more than a multi-label stratification. </p>",
          "rawMarkdown": "@bibek777 @hengck23 I am curious how this split is made too. In the previous competition steel and this competition, the split assignment has always been a puzzle. I think it does more than a multi-label stratification. "
        },
        {
          "id": 708607,
          "postDate": "2020-01-02T13:25:20.540Z",
          "content": "<p>A random split is what i am using now</p>",
          "rawMarkdown": "A random split is what i am using now",
          "votes": 1
        }
      ]
    },
    {
      "id": 760537,
      "postDate": "2020-03-01T12:41:06.157Z",
      "content": "<p>CamDrop: A New Explanation of Dropout and A Guided Regularization Method for Deep Neural Networks</p>\n\n<p><img src=\"https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/9f13cd4841fb050b345f3a56398871232ea5a58c/3-Figure1-1.png\" alt=\"\">\n<img src=\"https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/9f13cd4841fb050b345f3a56398871232ea5a58c/7-Table6-1.png\" alt=\"\"></p>",
      "rawMarkdown": "CamDrop: A New Explanation of Dropout and A Guided Regularization Method for Deep Neural Networks\n\n![](https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/9f13cd4841fb050b345f3a56398871232ea5a58c/3-Figure1-1.png)\n![](https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/9f13cd4841fb050b345f3a56398871232ea5a58c/7-Table6-1.png)",
      "votes": 2
    },
    {
      "id": 739298,
      "postDate": "2020-02-07T16:58:18.873Z",
      "content": "<p>imageBERT\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F679ae3e3b92efdb848b8a7ef3c4bc81c%2FSelection_067.png?generation=1581094696419061&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "imageBERT\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F679ae3e3b92efdb848b8a7ef3c4bc81c%2FSelection_067.png?generation=1581094696419061&amp;alt=media)\n",
      "votes": 2
    },
    {
      "id": 730971,
      "postDate": "2020-01-28T07:27:19.413Z",
      "content": "<p>to my surprise, the error are very subtle</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F910127fdfeb8365828a06159457427df%2FSelection_057.png?generation=1580196437091996&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "to my surprise, the error are very subtle\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F910127fdfeb8365828a06159457427df%2FSelection_057.png?generation=1580196437091996&amp;alt=media)\n",
      "votes": 2
    },
    {
      "id": 729042,
      "postDate": "2020-01-25T16:24:47.597Z",
      "content": "<p>error analysis\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcc5686df09164b9bded1f56fa081a095%2FSelection_124.png?generation=1579969484208260&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "error analysis\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcc5686df09164b9bded1f56fa081a095%2FSelection_124.png?generation=1579969484208260&amp;alt=media)\n",
      "votes": 2
    },
    {
      "id": 724604,
      "postDate": "2020-01-21T09:37:23.783Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F099b1bde79742f98fa45cba0c951d807%2FSelection_085.png?generation=1579599439688221&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F099b1bde79742f98fa45cba0c951d807%2FSelection_085.png?generation=1579599439688221&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 724607,
          "postDate": "2020-01-21T09:47:44.963Z",
          "content": "<p><a href=\"https://openreview.net/forum?id=rJMw747l_4\">https://openreview.net/forum?id=rJMw747l_4</a>\n\"To that end, we train ResNet-50 classifiers using either purely BigGAN images or mixtures of ImageNet and BigGAN images, and test on the ImageNet validation set.Our  preliminary  results  suggest  both a measured view of  state-of-the-art  GAN quality and highlight limitations of current metrics. Using only BigGAN images, we find that Top-1 and Top-5 error increased by 120% and 384%, respectively, and furthermore, adding more BigGAN data to the ImageNet training set at best only marginally improves classifier performance.\"</p>",
          "rawMarkdown": "https://openreview.net/forum?id=rJMw747l_4\n\"To that end, we train ResNet-50 classifiers using either purely BigGAN images or mixtures of ImageNet and BigGAN images, and test on the ImageNet validation set.Our  preliminary  results  suggest  both a measured view of  state-of-the-art  GAN quality and highlight limitations of current metrics. Using only BigGAN images, we find that Top-1 and Top-5 error increased by 120% and 384%, respectively, and furthermore, adding more BigGAN data to the ImageNet training set at best only marginally improves classifier performance.\"",
          "votes": 2
        }
      ]
    },
    {
      "id": 719277,
      "postDate": "2020-01-15T10:35:04.510Z",
      "content": "<p>just an idea\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbcaa1cc21e04038d9e9dafec16ae41e8%2FSelection_052.png?generation=1579084501734131&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "just an idea\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbcaa1cc21e04038d9e9dafec16ae41e8%2FSelection_052.png?generation=1579084501734131&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 719306,
          "postDate": "2020-01-15T11:14:16.930Z",
          "content": "<p>You are onto something. As an example আ should be classified as grapheme root: 3 with vowel diacritic: 0 and consonant diacritic: 0 whereas তা should be classified as grapheme root: 64 and vowel diacritic: 1. Though both have া, one has it as a part of the root and other don't. Similarly মু, ম,  both are same root but the former one has a vowel diacritic. Though the character itself has a shape  ু, it is important that the network can distinguish which is a part of root and which isn't.</p>",
          "rawMarkdown": "You are onto something. As an example আ should be classified as grapheme root: 3 with vowel diacritic: 0 and consonant diacritic: 0 whereas তা should be classified as grapheme root: 64 and vowel diacritic: 1. Though both have া, one has it as a part of the root and other don't. Similarly মু, ম,  both are same root but the former one has a vowel diacritic. Though the character itself has a shape  ু, it is important that the network can distinguish which is a part of root and which isn't.",
          "votes": 4
        }
      ]
    },
    {
      "id": 707041,
      "postDate": "2019-12-31T07:53:57.780Z",
      "content": "<p>Extremely happy to see you in the competition . Considering ,I have joined this competition to try out some basics and some new stuff . Your sharing will surely uplift the purpose of this competition and my own . One question though LSTM would normally help if you have a word written to get something contextual . The few paper I read it is used for that purpose . There could be some that I missed . Here the context that it can derive is if some grapheme combination does not occur in real life and if we can fine-tune the final result based on that ? </p>",
      "rawMarkdown": "Extremely happy to see you in the competition . Considering ,I have joined this competition to try out some basics and some new stuff . Your sharing will surely uplift the purpose of this competition and my own . One question though LSTM would normally help if you have a word written to get something contextual . The few paper I read it is used for that purpose . There could be some that I missed . Here the context that it can derive is if some grapheme combination does not occur in real life and if we can fine-tune the final result based on that ? ",
      "votes": 2,
      "replies": [
        {
          "id": 707078,
          "postDate": "2019-12-31T08:29:54.473Z",
          "content": "<p>since you are familiar with the language, human-in-the loop attention method may be suitable for you.\nyou can further create your own training samples (hard sample mining) as you mentioned in <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123002\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123002</a></p>\n\n<p><img src=\"https://d3i71xaburhd42.cloudfront.net/5393ccde31e38b1f278bc70af0c5b0a9cc758b96/3-Figure3-1.png\" alt=\"\"></p>\n\n<p>Embedding Human Knowledge in Deep Neural Network via Attention Map\n<a href=\"https://arxiv.org/abs/1905.03540\">https://arxiv.org/abs/1905.03540</a></p>\n\n<p><a href=\"https://github.com/machine-perception-robotics-group/attention_branch_network\">https://github.com/machine-perception-robotics-group/attention_branch_network</a></p>",
          "rawMarkdown": "since you are familiar with the language, human-in-the loop attention method may be suitable for you.\nyou can further create your own training samples (hard sample mining) as you mentioned in https://www.kaggle.com/c/bengaliai-cv19/discussion/123002\n\n![](https://d3i71xaburhd42.cloudfront.net/5393ccde31e38b1f278bc70af0c5b0a9cc758b96/3-Figure3-1.png)\n\nEmbedding Human Knowledge in Deep Neural Network via Attention Map\nhttps://arxiv.org/abs/1905.03540\n\nhttps://github.com/machine-perception-robotics-group/attention_branch_network",
          "votes": 10
        },
        {
          "id": 707248,
          "postDate": "2019-12-31T13:29:34.917Z",
          "content": "<p><img src=\"https://ars.els-cdn.com/content/image/1-s2.0-S0167865519302181-gr2.jpg\" alt=\"\"></p>",
          "rawMarkdown": "![](https://ars.els-cdn.com/content/image/1-s2.0-S0167865519302181-gr2.jpg)",
          "votes": 7
        }
      ]
    },
    {
      "id": 706324,
      "postDate": "2019-12-30T08:40:58.423Z",
      "content": "<p>How much did the swa help? any idea?</p>",
      "rawMarkdown": "How much did the swa help? any idea?",
      "votes": 2,
      "replies": [
        {
          "id": 717410,
          "postDate": "2020-01-13T04:46:23.400Z",
          "content": "<p>My score has been gone up 0.05</p>",
          "rawMarkdown": "My score has been gone up 0.05"
        },
        {
          "id": 718243,
          "postDate": "2020-01-14T08:04:35.903Z",
          "content": "<p>you mean 0.05 or 0.005? like 5 percent? that's a lot my god</p>",
          "rawMarkdown": "you mean 0.05 or 0.005? like 5 percent? that's a lot my god"
        }
      ]
    },
    {
      "id": 706924,
      "postDate": "2019-12-31T02:31:43.057Z",
      "content": "<p>Frog brother, See you again! Thx for your generous share; hope you get better rank</p>",
      "rawMarkdown": "Frog brother, See you again! Thx for your generous share; hope you get better rank",
      "votes": 1
    },
    {
      "id": 775962,
      "postDate": "2020-03-17T03:43:12.780Z",
      "content": "<p>Thanks for sharing the ideas! Helped a lot and inspired during the competition </p>",
      "rawMarkdown": "Thanks for sharing the ideas! Helped a lot and inspired during the competition "
    },
    {
      "id": 770235,
      "postDate": "2020-03-12T18:04:17.307Z",
      "content": "<p>getting the handwriting strokes</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=47CXR42_b2Y\">https://www.youtube.com/watch?v=47CXR42_b2Y</a></p>",
      "rawMarkdown": "getting the handwriting strokes\n\nhttps://www.youtube.com/watch?v=47CXR42_b2Y"
    },
    {
      "id": 770233,
      "postDate": "2020-03-12T18:02:41.877Z",
      "content": "<p>just wonder if anyone uses Tesseract?</p>",
      "rawMarkdown": "just wonder if anyone uses Tesseract?"
    },
    {
      "id": 770227,
      "postDate": "2020-03-12T17:51:34.303Z",
      "content": "<p>this is a common trick to train image of larger size (or deep network) with limited GPU resource:</p>\n\n<p>1) how to train large size:\n1. train image e.g. 360x360\n2. design fully convolutional net\n3. in training crop to smaller size, e.g. 224x224\n4. in final stage, fintunning at full image size (i.e. batch get smaller or freeze bottom layers)</p>\n\n<p>2) how to train very deep network\nthis is how vgg16 and googlenett did it when the gpu are small in the early days\n- train few layers first. then freeze these layers. add new trainable layers.\n  (or, implement deep network but alternate between freezing upper and bottom layers at training)</p>",
      "rawMarkdown": "this is a common trick to train image of larger size (or deep network) with limited GPU resource:\n\n1) how to train large size:\n1. train image e.g. 360x360\n2. design fully convolutional net\n3. in training crop to smaller size, e.g. 224x224\n4. in final stage, fintunning at full image size (i.e. batch get smaller or freeze bottom layers)\n\n\n2) how to train very deep network\nthis is how vgg16 and googlenett did it when the gpu are small in the early days\n- train few layers first. then freeze these layers. add new trainable layers.\n  (or, implement deep network but alternate between freezing upper and bottom layers at training)\n\n\n"
    },
    {
      "id": 737229,
      "postDate": "2020-02-05T04:23:30.307Z",
      "content": "<p>encoder based augmentation?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0cc99af813bda0e3f1ac4680dd8618ec%2FClipboard09.png?generation=1580876607555657&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "encoder based augmentation?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0cc99af813bda0e3f1ac4680dd8618ec%2FClipboard09.png?generation=1580876607555657&amp;alt=media)\n"
    },
    {
      "id": 737219,
      "postDate": "2020-02-05T04:06:46.347Z",
      "content": "<p>fine-grained classification\n<a href=\"https://arxiv.org/pdf/1903.06150v2.pdf\">https://arxiv.org/pdf/1903.06150v2.pdf</a></p>\n\n<p>Looking for the Devil in the Details: Learning Trilinear Attention Sampling\nNetwork for Fine-grained Image Recognition</p>\n\n<p><img src=\"https://paperswithcode.com/media/thumbnails/task/task-0000000719-ccaa52fd.jpg\" alt=\"\"></p>",
      "rawMarkdown": "fine-grained classification\nhttps://arxiv.org/pdf/1903.06150v2.pdf\n\nLooking for the Devil in the Details: Learning Trilinear Attention Sampling\nNetwork for Fine-grained Image Recognition\n\n![](https://paperswithcode.com/media/thumbnails/task/task-0000000719-ccaa52fd.jpg)\n"
    },
    {
      "id": 717619,
      "postDate": "2020-01-13T11:11:14.553Z",
      "content": "<p>Thanks for sharing this.</p>\n\n<p>You use a file \"grapheme_1295.csv\" for creating the balanced sampler. I have not seen any indication of what the content is or how the content is created. Could you please give some explanations?</p>",
      "rawMarkdown": "Thanks for sharing this.\n\nYou use a file \"grapheme_1295.csv\" for creating the balanced sampler. I have not seen any indication of what the content is or how the content is created. Could you please give some explanations?",
      "replies": [
        {
          "id": 718375,
          "postDate": "2020-01-14T10:38:05.637Z",
          "content": "<p>You can do something like this -<code>pd.DataFrame({'grapheme' :df['grapheme'].unique()}).to_csv('../grapheme_1295.csv')</code>  df - train.csv</p>",
          "rawMarkdown": "You can do something like this -` pd.DataFrame({'grapheme' :df['grapheme'].unique()}).to_csv('../grapheme_1295.csv')`  df - train.csv"
        },
        {
          "id": 718381,
          "postDate": "2020-01-14T10:40:48.763Z",
          "content": "<p>sample_submission.csv : just map grapheme symbol to numerical label, e.g.</p>\n\n<p>```\nlabel,grapheme\n0,ং\n1,ঃ\n2,অ\n3,অ্যা\n4,আ\n5,আঁ\n6,ই\n7,ইঁ\n8,ঈ\n9,উ\n10,উঁ\n11,ঊ</p>\n\n<p>```</p>",
          "rawMarkdown": "sample_submission.csv : just map grapheme symbol to numerical label, e.g.\n\n```\nlabel,grapheme\n0,ং\n1,ঃ\n2,অ\n3,অ্যা\n4,আ\n5,আঁ\n6,ই\n7,ইঁ\n8,ঈ\n9,উ\n10,উঁ\n11,ঊ\n\n```",
          "votes": 1
        },
        {
          "id": 719045,
          "postDate": "2020-01-15T04:07:56.087Z",
          "content": "<p>I have to thank you again, Heng. Your sharing always inspires me. <a href=\"/hengck23\">@hengck23</a> </p>",
          "rawMarkdown": "I have to thank you again, Heng. Your sharing always inspires me. @hengck23 "
        }
      ]
    },
    {
      "id": 717306,
      "postDate": "2020-01-13T02:15:26.040Z",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks buddy and best of luck!!!</p>",
      "rawMarkdown": "@hengck23 Thanks buddy and best of luck!!!"
    },
    {
      "id": 714390,
      "postDate": "2020-01-09T11:48:55.480Z",
      "content": "<p>Thanks for sharing, I will try to follow your step in this competition. Learned a lot from you </p>",
      "rawMarkdown": "Thanks for sharing, I will try to follow your step in this competition. Learned a lot from you "
    },
    {
      "id": 706664,
      "postDate": "2019-12-30T17:15:39.820Z",
      "content": "<p>I will stay tuned on your post.</p>",
      "rawMarkdown": "I will stay tuned on your post."
    },
    {
      "id": 761708,
      "postDate": "2020-03-02T21:19:23.920Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 759337,
      "postDate": "2020-02-28T22:49:45.423Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 722257,
      "postDate": "2020-01-18T10:21:43.973Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 706310,
      "postDate": "2019-12-30T08:30:20.873Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 771046,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-03-13T17:42:27.323000",
      "content": "<p>this is one of the kaggle competitions that i did the most number of experiments. I notice something interesting.</p>\n\n<p>for several experiments, it is important to apply the method at the early start of training. e.g. you get different results if you apply ohem on trained models (i.e. end stage of training) versus you apply it right at the start of training.</p>\n\n<p>i think that when data size is small, the number of good solutions (or generalized solutions) region is small. i.e. there are many local minimum. It is difficult to jump from one solution to another solution </p>",
      "votes": 8,
      "replies": [
        {
          "id": 771347,
          "author_name": "MadCoder",
          "author_url": "",
          "post_date": "2020-03-14T04:07:33.507000",
          "content": "<p>I do find out applying ohem at the end stage of training will save you from overfitting. It will shake a little bit at the beginning but it will increase your local cv at the end. It seems like applying ohem make it jumping out of its local minimum.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 756402,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-25T17:37:08.607000",
      "content": "<p>this is a reference  version for those who are struggling:</p>\n\n<p>version 20200224:\n- modified se-resnext50 (64x112 small input) with augmentation that works. please see readme file in folder. \n  This version can gives CV 0.991 and LB 0.980 after you optimized the augmentation hyper parameters  yourselves.</p>\n\n<ul>\n<li><p>it provides reference log file for you to compare the loss curve</p></li>\n<li><p>it has train/validation split that has gap of CV/LB 0.011 </p></li>\n<li><p>code is not complete, but should have enough details to reproduce the above mentioned results.\n(do not request me for missing files or functions)</p></li>\n</ul>",
      "votes": 7,
      "replies": [
        {
          "id": 756544,
          "author_name": "Andrey Zotov",
          "author_url": "",
          "post_date": "2020-02-25T20:23:17.407000",
          "content": "<p>Thanks for sharing and for the best thread in the discussion!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756562,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-25T20:56:28.917000",
          "content": "<p>Hi Heng Thanks for the update. May I ask if the 64x112 is a simple downsize from the original 137x236 image or did you do some post processing?</p>\n\n<p>And may I ask how long have you trained the model to reach cv ~.99. From that log file it seems it's been training for 500 epochs?!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756712,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-26T01:43:17.457000",
          "content": "<p>please refer to the readme file</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 757358,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-26T17:09:16.117000",
          "content": "<p>how to train a modified resenext model with pretrained imagenet model</p>\n\n<ol>\n<li><p>use the unmodified version resenext model first. initialised with pretrained imagenet model and train as usual.</p></li>\n<li><p>assume resenext = block0, block1 ... block4. we modify block0.</p></li>\n<li><p>load the previously trained model except for the the modified part (i.e block0).</p></li>\n<li><p>freeze all modified part (including the classifier head). Train and the gradient will only back-propagated at the modified block0.</p></li>\n<li><p>when the accuracy/loss for the modified model is about the same as the previously unmodified one, say 90% of the results, unfreeze all  model and proceed training as usual.</p></li>\n</ol>\n\n<p>this is how i do network surgery. apply this when there is structural change of the model, e.g. add new convolution layers, reduce number of channels, change kernel size, ...</p>\n\n<hr>\n\n<p>if the old pretrained imagenet model can be loaded after modification (e.g. simple modification like changing strike), there is no need to freeze and unfreeze. In such minor modification, just train as usual.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 758076,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-27T12:16:27.667000",
          "content": "<p>purely trained from scratch \"without surgery above\".\nno augmentation is used at all!</p>\n\n<p>local cv = 0.972 (at rate 0.05)\nlocal cv = 0.972 (at rate 0.005)</p>\n\n<p>training loss is almost zero, so there is no point to train further. expected LB should be around 0.960 (without augmentation)</p>\n\n<p>see attached file.\n```\ndef valid_augment(image, label, infor):\n    image = to_64x112(image)\n    return image, label, infor</p>\n\n<p>def train_augment(image, label, infor):\n    original = image.copy()\n    original = to_64x112(original)\n    image = to_64x112(image)\n    return original, image, label, infor</p>\n\n<p>def train_batch_augment(original, input, onehot):\n    return input, onehot</p>\n\n<p>0.05000   9.0   15.6 | 0.972 : 0.957 0.989 0.982  0.955 | 0.20, 0.06, 0.06, 0.210 : 0.96, 0.99, 0.99, 0.956 | 0.00, 0.00, 0.00, 0.009 | 1 hr 38 min\n0.05000  10.0   17.3 | 0.972 : 0.957 0.990 0.986  0.955 | 0.21, 0.07, 0.07, 0.217 : 0.96, 0.99, 0.99, 0.956 | 0.00, 0.00, 0.00, 0.007 | 1 hr 49 min\n0.05000  11.0   19.0 | 0.971 : 0.957 0.989 0.982  0.955 | 0.21, 0.06, 0.07, 0.214 : 0.96, 0.99, 0.99, 0.956 | 0.01, 0.00, 0.00, 0.008 | 2 hr 00 min\n0.05000  12.0*  20.8 | 0.970 : 0.955 0.990 0.982  0.953 | 0.22, 0.06, 0.07, 0.224 : 0.96, 0.99, 0.99, 0.954 | 0.00, 0.00, 0.00, 0.007 | 2 hr 11 min\n```\nnote: some kaggler reported about to get LB 0.974 without augmentation (which i think 0.984 for local CV). hence i need to investigate better model or regularization (like dropout, shakedrop, etc) or better optimiser</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 758342,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-27T16:49:19.633000",
          "content": "<p>additional information:</p>\n\n<p>modified se-resnext50 (32x56 mini input) with same augmentation that works  gives CV 0.987\nthis may be useful for auto-augmentation parameters search?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 759575,
          "author_name": "thefatcat",
          "author_url": "",
          "post_date": "2020-02-29T07:58:51.023000",
          "content": "<p>Thanks for your sharing,but I have a question.\nIn your train_rand_aug2c.py,there is a func named INVERSE_COMPOSE,but you didn't assign it before using，I guess it is used to add information for first three predict form firth .\nCould you explain its meaning in detail? Thanks！</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759588,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-29T08:24:02.713000",
          "content": "<p>we know that each of the 1295 grapheme is made up of grapheme=root, vowel, consonant</p>\n\n<p>so if you have predicted the 1295 class, simply  decode (root, vowel, consonant) = grapheme</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759960,
          "author_name": "thefatcat",
          "author_url": "",
          "post_date": "2020-02-29T16:56:38.817000",
          "content": "<p>Oh,yes.I understand your idea.</p>\n\n<p>But I got another question when I use your code to train the model .Most of the time my GPU utilization ratio is 0 and I thought the reason for this is dataloder's speed limitation.When we load the data by dataloader,we do some augmentation,and before send data to the net,maybe do some mixup or cutout.These process are not handled by GPU but CPU.</p>\n\n<p>This problem made my training process very slow,but in your log.train,it shows the speed is not so slow,so I want to ask whether you use some tricks to speed up.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759969,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-29T17:13:32.273000",
          "content": "<p>df['xxx'].values should be done once at initialisation and not everytime at getitem()</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759981,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-29T17:27:09.323000",
          "content": "<p>```</p>\n\n<p>class KaggleDataset(Dataset):\n    def <strong>init</strong>(self, split, mode, csv, parquet, augment=None):\n        global TRAIN_PARQUET\n        ...\n        self.df_values = df.values</p>\n\n<pre><code>def __getitem__(self, index): \n    i, image_id, grapheme_root, vowel_diacritic, consonant_diacritic, grapheme, grapheme_symbol  =  self.df_values[index]\n</code></pre>\n\n<p>```</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 760463,
          "author_name": "thefatcat",
          "author_url": "",
          "post_date": "2020-03-01T10:33:43.803000",
          "content": "<p>Thank you,I will have a try.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764147,
          "author_name": "PY",
          "author_url": "",
          "post_date": "2020-03-05T07:19:47",
          "content": "<p>Thanks for sharing, very helpful! Quick question, when I'm using your model <code>ResNext50</code>, it seems that the model takes more than 11GB of GPU memory. Did you encounter similar issue? </p>\n\n<p>Or does this model only work when you freeze part of the model and only train its head?</p>\n\n<p>Thanks. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765755,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-03-07T03:43:19.483000",
          "content": "<p>I am still trying to figure out how to write inverse_compose function.  Any idea ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766504,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-03-08T09:09:05.520000",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> </p>\n\n<p>```\ndef run_prepare_data4():\n    df_grapheme = pd.read_csv(DATA_DIR+'/grapheme_1295.csv')\n    df_root = pd.read_csv(DATA_DIR+'/root_168.csv')\n    df_vowel = pd.read_csv(DATA_DIR+'/vowel_11.csv')\n    df_consonant = pd.read_csv(DATA_DIR+'/consonant_7.csv')</p>\n\n<pre><code>grapheme_map = dict(df_grapheme[['grapheme','label']].values)\nroot_map = dict(df_root[['root','label']].values)\nvowel_map = dict(df_vowel[['vowel','label']].values)\nconsonant_map = dict(df_consonant[['consonant','label']].values)\n\ngrapheme_inverse_map = dict(df_grapheme[['label','grapheme']].values)\nroot_inverse_map = dict(df_root[['label','root']].values)\nvowel_inverse_map = dict(df_vowel[['label','vowel']].values)\nconsonant_inverse_map = dict(df_consonant[['label','consonant',]].values)\n\n#----\ndf_train = pd.read_csv(DATA_DIR+'/train.csv')\ndf_train['grapheme']=df_train['grapheme'].map(grapheme_map)\nassert(df_train['grapheme'].isnull().values.any() == False)\n\ncompose = {}\ninverse_compose = {}\ngb = df_train.groupby(['grapheme',])\nfor i in range(1295):\n    gp = gb.get_group(i) #.index\n\n    gp = gp.drop_duplicates(subset=['grapheme_root', 'vowel_diacritic', 'consonant_diacritic'])\n    assert(len(gp)==1)\n\n    g = i\n    r,v,c = gp[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']].values[0]\n\n    inverse_compose[g]= (r,v,c)\n    compose[grapheme_inverse_map[g]]= (\n        root_inverse_map[r],\n        vowel_inverse_map[v],\n        consonant_inverse_map[c]\n    )\n\nprint(inverse_compose)\nprint(compose)\n</code></pre>\n\n<p>```</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 766512,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-03-08T09:21:13.480000",
          "content": "<p>Thanks so much Heng :) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 737223,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-05T04:15:29.610000",
      "content": "<p>i sometime have discussions with friends on their phd research work. recently, they give me a few pointers on network design: </p>\n\n<ol>\n<li><p>it is possible that a single batch norm cannot normalized the feature well for multi-task. different task needs different normalization.  Care needs to be taken at designing the head</p></li>\n<li><p>you can force the network to learn high frequency features (e.g.  edge, etc) over texture  by adding loss related to high frequency feature extraction , together with the usual cross entropy classification loss.</p></li>\n</ol>",
      "votes": 7,
      "replies": [
        {
          "id": 737512,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-05T12:57:30.470000",
          "content": "<p>related:\nusing Adversarial Examples to expand dataset</p>\n\n<p>\"With an enhanced EfficientNet-B8, our method achieves the state-of-the-art 85.5% ImageNet\ntop-1 accuracy without extra data. This result even surpasses the best model in [20] which is trained with 3.5B Instagram images (∼3000× more than ImageNet) and ∼9.4× more parameters.\"</p>\n\n<p>Adversarial Examples Improve Image Recognition\n<a href=\"https://arxiv.org/abs/1911.09665\">https://arxiv.org/abs/1911.09665</a>\n<a href=\"https://www.youtube.com/watch?v=KTCztkNJm50\">https://www.youtube.com/watch?v=KTCztkNJm50</a></p>\n\n<p><img src=\"https://i.redd.it/8rm53y9puf141.png\" alt=\"\"> <img src=\"https://neurohive.io/wp-content/uploads/2019/11/Screenshot-from-2019-11-26-23-34-49-570x419.png\" alt=\"\"></p>\n\n<p>see also:\nCompounding the Performance Improvements of Assembled Techniques in a Convolutional Neural Network <a href=\"https://arxiv.org/pdf/2001.06268.pdf\">https://arxiv.org/pdf/2001.06268.pdf</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 723886,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-20T15:16:12.777000",
      "content": "<p>magic net</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd43e2775c784672d47d6b37cc582cb90%2FSelection_073.png?generation=1579533369553774&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd5ff591bb2d20c7c87753e6c2170189b%2FSelection_079.png?generation=1579594728840307&amp;alt=media\" alt=\"\"></p>",
      "votes": 8,
      "replies": [
        {
          "id": 724702,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-21T11:56:12.537000",
          "content": "<p>single model, single fold performance:\n```\nLB = 0.9755\nCV = 0.983629\n               grapheme_root  0.977137\n             vowel_diacritic  0.993395\n         consonant_diacritic  0.986848 </p>\n\n<p>model: SE-Resnext50 with stride replaced with max pooling + SWA\ninput: 137x236\naugmentation: mixup + cut-fade</p>\n\n<p>```</p>\n\n<p>after more tunning:</p>\n\n<p>```\nLB = 0.9760\nCV = 0.984155\n               grapheme_root  0.978920\n             vowel_diacritic  0.992851\n         consonant_diacritic  0.985929</p>\n\n<p>```</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 726108,
          "author_name": "Mario Parreño Lara",
          "author_url": "",
          "post_date": "2020-01-22T20:27:47.137000",
          "content": "<p>Hi Heng! I am trying to automate the 'magic' net architecture. First I replaced de stride 2 to 1, saving the conv 'path':</p>\n\n<p><code>\nmodules = {}\nfor name, module in model.named_modules():\n    if(isinstance(module, nn.Conv2d)):\n        stride = module.stride\n        if stride == (2, 2) or stride == 2:\n            module.stride = (1,1)\n            modules[name] = module\n        elif stride == 2:\n            module.stride = 1\n            modules[name] = module\n</code></p>\n\n<p>Next I try to insert the maxpool as follows:</p>\n\n<p>```\nfor name in modules:\n    parent_module = model\n    objs = name.split(\".\")\n    if len(objs) == 1:\n        #model.<strong>setattr</strong>(name, modules[name])\n        model.<strong>setattr</strong>(\"magicMaxPool\", nn.MaxPool2d(kernel_size=2, stride=2))\n        continue</p>\n\n<pre><code>for obj in objs[:-1]:\n    parent_module = parent_module.__getattr__(obj)\n\nparent_module.__setattr__(\"magicMaxPool\", nn.MaxPool2d(kernel_size=2, stride=2))\n</code></pre>\n\n<p>```</p>\n\n<p>If I inspect the net/modules, it appears the maxpool but when I do the inference stride works properly but maxpool is not used (vector size before global avg pool is too big).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 729048,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-25T16:46:53.873000",
          "content": "<p>related: <a href=\"http://openaccess.thecvf.com/content_cvpr_2017/papers/Zhai_S3Pool_Pooling_With_CVPR_2017_paper.pdf\">http://openaccess.thecvf.com/content_cvpr_2017/papers/Zhai_S3Pool_Pooling_With_CVPR_2017_paper.pdf</a></p>\n\n<p>S3Pool: Pooling with Stochastic Spatial Sampling</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 735731,
          "author_name": "makogarei",
          "author_url": "",
          "post_date": "2020-02-03T11:53:18.557000",
          "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> \nI always look at the code you write and I am very grateful.\nI try to training Resnext50 with magic net,But cv value did not go above 0.96. Did you do anything other than change the layer?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 707588,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-01T05:35:10.807000",
      "content": "<p>after a series of experiments, how are my conclusion:</p>\n\n<ol>\n<li>final first LB ranking: 0.99+ after 3 months</li>\n<li>shakeup: +/- 0.002 (should be quite small)</li>\n</ol>\n\n<p>baseline results:\n a. 0.9675~0.9715\n  - standard method (joint 3 class) and ensemble standard model (e.g. densenet, resnet, efficientnet, etc) \n  - to maximize performance, search for best input size (e.g. original size, enlarged size like 224x224,256x256, or even larger), try better augmentation, train longer with better learning rate (swa, cyclic annealing, ...), etc\n  - size, scale normalization</p>\n\n<p>b. 0.9700~0.9800\n... to be updated ... (better formulation, loss etc?)</p>\n\n<p>c. 0.9800~0.9900\n... to be updated ... (some way to create more data?)</p>\n\n<p>don't forget  Google Cloud AutoML Vision as part of solution !!!!</p>",
      "votes": 8,
      "replies": [
        {
          "id": 708848,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-02T19:00:30.557000",
          "content": "<p>Thanks for the pointers <a href=\"/hengck23\">@hengck23</a>  . I started playing with augmentations . I have got the cutout working here . </p>\n\n<p><a href=\"https://www.kaggle.com/phoenix9032/pytorch-efficientnet-starter-code\">https://www.kaggle.com/phoenix9032/pytorch-efficientnet-starter-code</a></p>\n\n<p>Now , need to see how MixUp and RICAP and other thing can work out along with the training for multi-heads . Need to make some modification from the original implementation.</p>\n\n<p>We can try dual cutout as well. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 719435,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-15T14:01:09.920000",
          "content": "<p>some important (?) experimental results</p>\n\n<ul>\n<li><p>sampling affects results (e.g. balance sampling of \"root\" or \"grapheme\" affects \"root\", \"constant\", \"vowel\" differently)</p></li>\n<li><p>augmentation affects results (e.g. some drop in accuracy when scaling and rotation is used)</p></li>\n<li><p>arcface and large-margin like softamx does improve results but cannot be used with mixup?</p></li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 721090,
          "author_name": "FP",
          "author_url": "",
          "post_date": "2020-01-17T03:20:43.097000",
          "content": "<p>Hi Heng <a href=\"/hengck23\">@hengck23</a> (and other Kagglers), I see in your latest starting kit that you added a fourth target of the 1295 unique classes of Bengali letters as provided in the training dataset. If the testing dataset (which is not accessible to us) has letters other than those 1295 classes, would the training approach be prone to overfitting? Please correct me if I am wrong. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721132,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-17T04:53:19.677000",
          "content": "<p>you can use low weight like 0.1 for the fourth class, you can also use 1295+1 (background class for none of the above). to create background class, selectively flip the original train samples</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 758757,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-28T05:38:01.697000",
      "content": "<p>good for augmentation?</p>\n\n<p><img src=\"https://raw.githubusercontent.com/warbean/tps_stn_pytorch/master/demo/top_1.gif\" alt=\"\"></p>\n\n<p><a href=\"https://github.com/WarBean/tps_stn_pytorch\">https://github.com/WarBean/tps_stn_pytorch</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 759014,
          "author_name": "Soonhwan Kwon",
          "author_url": "",
          "post_date": "2020-02-28T13:02:09.523000",
          "content": "<p>very interesting augmemtation!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 724848,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-21T14:51:04.107000",
      "content": "<p>in google doddle, you can add and subtract drawing. I wonder if this can be done in bengali grapheme? e.g. add/remove or replace vowel, etc\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F249a956c438d0dd9ef6f618882dde10a%2FSelection_087.png?generation=1579618261743829&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 723868,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-20T14:54:31.900000",
      "content": "<p>new augmentation</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb35fa0d1187ed81acb31b38fedbc6828%2FSelection_069.png?generation=1579532069051469&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 724274,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-21T01:59:26.787000",
          "content": "<p>related: augmentation on feature map:</p>\n\n<ol>\n<li>choose some random feature map at training</li>\n<li>choose same max value</li>\n<li>modify: new value = alpha * old value, where alpha is between 0 to 1 (preferably near 0.5)</li>\n</ol>\n\n<p>this attenuate feature values, preventing it from over dominating and lead to over fitting </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 775716,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-03-17T00:23:49.243000",
      "content": "<p>private lb of coded posted here:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fda66de721f1d8955b9925e3642f5a0f0%2FSelection_074.png?generation=1584404626455936&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 775910,
          "author_name": "hawkey",
          "author_url": "",
          "post_date": "2020-03-17T02:45:02.400000",
          "content": "<p>Amazing results but sorry to see they are not selected...This post is really helpful and I have learned a lot from your code for the first competition I decided to take part in seriously. Thank you!\nBy the way, did you forget to select final submission?😂  Or just because of the shake up?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 765573,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-03-06T20:08:44.370000",
      "content": "<p>there is one extra supervision signal.</p>\n\n<p><code>\n list('ক্ট্রো')\n['ক', '্', 'ট', '্', 'র', 'ো']\n</code>\na list command breaks the graheme into more component. you can use this the measure the amount of common parts (distance) between 2 graphemes </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 761546,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-03-02T16:38:15.470000",
      "content": "<p>if this model is not correlated to image model, it is good for ensemble.\nthis work enables input of very large image</p>\n\n<p>Learning in the Frequency Domain\nKai Xu, Minghai Qin, Fei Sun, Yuhao Wang, Yen-kuang Chen, Fengbo Ren</p>\n\n<p><img src=\"https://storage.googleapis.com/groundai-web-prod/media/users/user_12457/project_409932/images/x2.png\" alt=\"\"></p>\n\n<p><img src=\"https://storage.googleapis.com/groundai-web-prod/media/users/user_12457/project_409932/images/x4.png\" alt=\"\"></p>\n\n<p>```</p>\n\n<p>ResNet-50   #Channels   Size Per Channel    Top-1   Top-5   Normalized Input Size <br>\nRGB             3               224x224                   75.780    92.650  1.0 <br>\nDCT-24 (ours)    24             56x56                     77.196    93.504  0.5      </p>\n\n<p>```</p>",
      "votes": 3,
      "replies": [
        {
          "id": 763297,
          "author_name": "momi64",
          "author_url": "",
          "post_date": "2020-03-04T10:37:02.583000",
          "content": "<p>Thanks for sharing an interesting paper.\nThe code is available below.</p>\n\n<p><a href=\"https://github.com/calmevtime1990/supp/tree/master/classification\">https://github.com/calmevtime1990/supp/tree/master/classification</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 728658,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-25T03:28:10.440000",
      "content": "<p>top-1 and top-2 recall:</p>\n\n<p>```\nlocal validation results</p>\n\n<p>avgerage recall (top-1): 0.982253\n               grapheme_root  0.974616\n             vowel_diacritic  0.992676\n         consonant_diacritic  0.987103</p>\n\n<p>avgerage recall (top-2) : 0.995918\n               grapheme_root  0.993451\n             vowel_diacritic  0.998653\n         consonant_diacritic  0.998116</p>\n\n<p>```</p>\n\n<p>if you make a mistake, the correct results is probably the top-2. if you can think of a good way to post-process, you can improve your score</p>",
      "votes": 3,
      "replies": [
        {
          "id": 729149,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-01-25T20:10:12.520000",
          "content": "<p>this screams for an ensemble of models ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 757812,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-27T05:58:26.333000",
      "content": "<p>New paper today\nOn Feature Normalization and Data Augmentation</p>\n\n<p><a href=\"https://arxiv.org/pdf/2002.11102.pdf\">https://arxiv.org/pdf/2002.11102.pdf</a></p>",
      "votes": 4,
      "replies": [
        {
          "id": 757817,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-27T06:09:36.200000",
          "content": "<p>arXiv:2002.11022 (cross-list from cs.LG) [pdf, other]\nBeyond Dropout: Feature Map Distortion to Regularize Deep Neural Networks\nYehui Tang, Yunhe Wang, Yixing Xu, Boxin Shi, Chao Xu, Chunjing Xu, Chang Xu\nSubjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 706574,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2019-12-30T15:27:50.050000",
      "content": "<p>What's the meaning of “without bn refinement”? </p>",
      "votes": 3,
      "replies": [
        {
          "id": 706611,
          "author_name": "Uday Kamal",
          "author_url": "",
          "post_date": "2019-12-30T16:07:09.213000",
          "content": "<p>I think Heng is talking about the recently proposed 'Full Normalization' technique, a better alternative (or you can say modification) to the conventional Batch Normalization. The idea is described in the following paper:\n<a href=\"https://arxiv.org/abs/1810.06177\">https://arxiv.org/abs/1810.06177</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 706618,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2019-12-30T16:11:02.760000",
          "content": "<p><a href=\"https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/\">https://pytorch.org/blog/stochastic-weight-averaging-in-pytorch/</a></p>\n\n<p>```\nBATCH NORMALIZATION\nOne important detail to keep in mind is batch normalization. Batch normalization layers compute running statistics of activations during training. Note that the SWA averages of the weights are never used to make predictions during training, and so the batch normalization layers do not have the activation statistics computed after you reset the weights of your model with opt.swap_swa_sgd(). To compute the activation statistics you can just make a forward pass on your training data using the SWA model once the training is finished. In the SWA class we provide a helper function opt.bn_update(train_loader, model). It updates the activation statistics for every batch normalization layer in the model by making a forward pass on the train_loader data loader. You only need to call this function once in the end of training.</p>\n\n<p>```</p>\n\n<p>the above step is not implemented</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 706626,
          "author_name": "Uday Kamal",
          "author_url": "",
          "post_date": "2019-12-30T16:19:48.630000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  Thanks for your clarification! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 706492,
      "author_name": "Wang Xinliang",
      "author_url": "",
      "post_date": "2019-12-30T13:15:02.323000",
      "content": "<p>What's the meaning of swa ?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 706607,
          "author_name": "Uday Kamal",
          "author_url": "",
          "post_date": "2019-12-30T16:03:32.357000",
          "content": "<p>Stochastic Weight Averaging, you can read this paper for more details:\n<a href=\"https://arxiv.org/abs/1803.05407\">https://arxiv.org/abs/1803.05407</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 707056,
          "author_name": "Wang Xinliang",
          "author_url": "",
          "post_date": "2019-12-31T08:05:01.543000",
          "content": "<p>thx</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 728668,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-25T04:01:19.887000",
      "content": "<p>multi-task and task dependency</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fefb21bf6c5f708e84efb92a2a610f24f%2FClipboard05.png?generation=1579924874232793&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0b65fbdbc95cf5730ca99eff92e7bf6f%2FClipboard04.png?generation=1579924877627631&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 740985,
          "author_name": "Tobepellucid",
          "author_url": "",
          "post_date": "2020-02-10T03:46:02.093000",
          "content": "<p>Im studying the correlation among tasks. What's this paper? Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 726161,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-22T21:58:56.863000",
      "content": "<p>loss formulation:</p>\n\n<p>some task (root, constant, or vowel) are more confidence than others.\nsimilarly, some classes are easier within a class.</p>\n\n<p>e.g. given task 1 has class A,B,C,D,E and task 2 has class a,b,c and if we know that some class cannot co-exist, it may give better results if we apply the rules as follow:\n```\n1. naive method:\ntreat task 1 as 5 classes  problem and task 2 as 3 classes. ignore any co-occurrence prior knowledge</p>\n\n<ol>\n<li>use prior knowledge\nif input is a, then it can only be A,B,C\nif input is b, then it can only be A,B,E\nif input is c, then it can only be D,E</li>\n</ol>\n\n<p>```\nif you know the language, you may apply such rules</p>",
      "votes": 4,
      "replies": [
        {
          "id": 726925,
          "author_name": "Dhananjay Raut",
          "author_url": "",
          "post_date": "2020-01-23T10:16:45.180000",
          "content": "<p>or we could change the architecture of model's head such that at the time of predicting task 2 we feed the probabilities(softmax) for task 1 predicted by our model.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 716656,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-12T04:17:58.407000",
      "content": "<p>version 20200111:\nse-resnext50 + balanced sampler + mixup (please refer to readme.ppt at the google drive)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F97673c14b01d4aff29fc1d688b036267%2FSelection_069.png?generation=1578830084131882&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 717189,
          "author_name": "FP",
          "author_url": "",
          "post_date": "2020-01-12T21:04:28.410000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> thanks for your sharing again. Do you have the updated verison of <code>compute_kaggle_metric</code> in your <code>kaggle.py</code>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717592,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-13T10:50:10.053000",
          "content": "<p>```\ndef compute_kaggle_metric(probability, truth):</p>\n\n<pre><code>def compute_recall(probability,truth):\n    num_class = probability.shape[-1]\n    y = probability.argmax(-1)\n    t = truth\n    correct = y==t\n\n    recall = np.zeros(num_class)\n    for c in range(num_class):\n        e = correct[t==c]\n        if len(e)&amp;gt;0:\n            recall[c]=e.mean()\n    return recall\n\ncomponet = []\nrecall   = []\nfor p,t in zip(probability,truth):\n    r = compute_recall(p,t)\n    recall.append(r)\n    componet.append(r.mean())\n\naverage = np.average(componet, weights=[2,1,1,0])\nreturn average, componet, recall\n</code></pre>\n\n<p>```</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 706712,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-12-30T18:06:51.830000",
      "content": "<p>there are 168 grapheme_root, 11 vowel_diacritic, 7 consonant_diacritic and 1295 grapheme classes!\nyou can build a fourth classifier for 1295 grapheme classes and decode the component from it</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbeafecfbcd56e0f8c5ee4ed6d98bd2b3%2FSelection_088.png?generation=1577729205397481&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0e6843ab6fef11c4b7b478c45d0f07f2%2FSelection_087.png?generation=1577729209517703&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 706932,
          "author_name": "Hanjoon Choe",
          "author_url": "",
          "post_date": "2019-12-31T03:04:58.270000",
          "content": "<p>Find a way to balance per class accuracy. &lt;- I thought this approach too and tried a bit although I don’t have convincing reason why it is helpful. Do you have it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717364,
          "author_name": "Yangfan",
          "author_url": "",
          "post_date": "2020-01-13T03:58:19.630000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759126,
          "author_name": "thefatcat",
          "author_url": "",
          "post_date": "2020-02-28T16:12:54.400000",
          "content": "<p>I was curious about how to build a fourth classifier for 1295 grapheme classes,is that just groupby grapheme and label it from 1 to 1295?\nHope you could clarify,thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759134,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-28T16:19:22.130000",
          "content": "<p><a href=\"/thefatcat\">@thefatcat</a> </p>\n\n<p>yes, you are correct</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 706395,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-12-30T10:31:29.073000",
      "content": "<p>augmentation in the release 'version 20191230'</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F13e81e022c9b8ebf067517e15d62e938%2FSelection_081.png?generation=1577701870877750&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 706738,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2019-12-30T18:55:33.683000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  Hi. Can u tell what augmentations u used ?  Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 716020,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2020-01-11T06:47:03.763000",
      "content": "<p>I published a <a href=\"https://www.kaggle.com/bibek777/heng-s-starter-training-kernel\">kernel</a> based on this so that people can train it using kaggle GPUs</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 760615,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-03-01T14:25:04.707000",
      "content": "<p>yet another paper:\nMaxUp: A Simple Way to Improve Generalization of Neural Network Training\n```</p>\n\n<p>Method  Top-1 error Top-5 error\nVanilla (He et al., 2016a)  76.3    -\nDropout (Srivastava et al., 2014)   76.8    93.4\nDropPath (Larsson et al., 2017) 77.1    93.5\nManifold Mixup (Verma et al., 2019) 77.5    93.8\nAutoAugment (Cubuk et al., 2019a)   77.6    93.8\nMixup (Zhang et al., 2018)  77.9    93.9\nDropBlock (Ghiasi et al., 2018) 78.3    94.1\nCutMix (Yun et al., 2019)   78.6    94.0\nMaxUp+CutMix    78.9    94.2</p>\n\n<p>Table 1: Summary of top1 and top5 accuracies on the validation set of ImageNet for ResNet-50.\n```</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 758761,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-28T05:43:32.973000",
      "content": "<p>it is possible to treat it as a detection problem\n<img src=\"https://raw.githubusercontent.com/RubanSeven/CRAFT_keras/master/images/CRAFT%E9%AB%98%E6%96%AF%E7%83%AD%E5%9B%BE.png\" alt=\"\"></p>\n\n<p><img src=\"https://pythonawesome.com/content/images/2019/06/CRAFT.jpg\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 759015,
          "author_name": "Soonhwan Kwon",
          "author_url": "",
          "post_date": "2020-02-28T13:02:58.843000",
          "content": "<p>Didn't expect to see CRAFT here</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759082,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-28T14:49:24.743000",
          "content": "<p>just an idea\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F2e6d456f54d997d4cfabfd1ad5a6af41%2FSelection_082.png?generation=1582901362622740&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 758090,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-27T12:35:38.497000",
      "content": "<p><a href=\"https://github.com/sujatasaini/Kuzushiji-DropBlock\">https://github.com/sujatasaini/Kuzushiji-DropBlock</a></p>\n\n<p>|Models                           | MNIST | Fashion-MNIST | Kuzushiji-MNIST | Kuzushiji-49 |\n|---------------------------------|-------|---------------|-----------------|--------------|\n|DCNN-DropBlock     | <strong>99.47%</strong> | <strong>93.40%</strong> | <strong>97.66%</strong> | <strong>95.67%</strong> |\n|DCNN-Dropout                         | 97.99% | 85.47% | 86.43% | 95.34% |\n|DCNN-Spatial-Dropout          | 97.17% | 84.44% |  81.08% | 58.18 |</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 751629,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-02-20T10:55:19.673000",
      "content": "<p>another augmentation that works for me:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F5800610b7d3b4f3f03d8e9675b889030%2FSelection_085.png?generation=1582196108850096&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"http://yann.lecun.com/ex/images/invar.png\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 751631,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-02-20T10:57:03.813000",
          "content": "<p>```\ndef train_augment(image, label, infor):\n    original = image.copy()\n    if 1:\n        operation = [\n            lambda image : do_random_erode(image),\n            lambda image : do_random_dilate(image),\n            lambda image : do_random_grid_distortion(image, distort=0.10, num_step = 5),\n            lambda image : do_random_custom_distortion1(image, distort=0.10),\n            lambda image : do_random_crop_rotate_rescale(image, mode={'rotate':10}),\n            lambda image : do_random_block_fade(image, size=[0.1, 0.9], alpha=(0.1,0.5)),\n            lambda image : do_random_line(image),\n        ]\n        num_op = np.random.choice(2)\n        if num_op&gt;0:\n            for op in np.random.choice(operation,num_op):\n                image = op(image)\n        image = do_random_pad_crop(image, 3)</p>\n\n<pre><code>#----------------------------------------------------\nif 1:\n    operation = [\n        lambda image : image,\n        lambda image : do_random_salt_pepper(image),\n        lambda image : do_random_noise(image),\n    ]\n    for op in np.random.choice(operation,1):\n        image = op(image)\n\n#----------------------------------------------------\nreturn original, image, label, infor\n</code></pre>\n\n<p>def train_batch_augment(original, input, onehot):\n    operation = [\n        lambda input, onehot : (input, onehot),\n        lambda input, onehot : do_random_batch_cutout(input, onehot, fill=0),\n        lambda input, onehot : do_random_batch_mixup(original, onehot),\n    ]\n    op = np.random.choice(operation)\n    with torch.no_grad():\n        input, onehot = op(input, onehot)</p>\n\n<pre><code>return input, onehot\n</code></pre>\n\n<p>```</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 761929,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-03T03:36:04.637000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 737055,
      "author_name": "Rohit Agarwal",
      "author_url": "",
      "post_date": "2020-02-04T21:26:15.080000",
      "content": "<p>This discussion forum is a gold-mine of ideas. Thanks a lot <a href=\"/hengck23\">@hengck23</a> for initiating this and for sharing such wonderful ideas. </p>\n\n<p>Lots to do, lots of learn, not enough time.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 733502,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-31T08:15:44.307000",
      "content": "<p>don't forget that each of has has addition 36 unlabelled training images at the test parquet files</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 733430,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-31T05:30:36.087000",
      "content": "<p>some tricks from ICDAR compeitions</p>\n\n<p><a href=\"https://gateway.newton.ac.uk/sites/default/files/asset/doc/1811/20181127-Large%20Scale%20and%20Unconstrained%20OHCCR%20Based%20on%20Deep%20Learning%20and%20Path%20Signature.pdf\">https://gateway.newton.ac.uk/sites/default/files/asset/doc/1811/20181127-Large%20Scale%20and%20Unconstrained%20OHCCR%20Based%20on%20Deep%20Learning%20and%20Path%20Signature.pdf</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 733436,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-31T05:52:45.607000",
          "content": "<p>handwriting trajectory recovery\n<img src=\"https://encrypted-tbn0.gstatic.com/images?q=tbn%3AANd9GcQ_0pEXIOBFOypFTak5Pq9drW8LvXBx_W0EdDOoAC4jDTyYO3mf\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 733846,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-31T15:20:33.237000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F1001950b2fa0efdac57746685d495781%2FSelection_066.png?generation=1580484022666591&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://arxiv.org/pdf/1606.05763.pdf\">https://arxiv.org/pdf/1606.05763.pdf</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 733850,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-31T15:24:05.603000",
          "content": "<p>dropsample --- remove outlier\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F631f809daaa2cdbbbe2938d1efd1b8ed%2FSelection_067.png?generation=1580484242967773&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 733853,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-31T15:30:55.200000",
          "content": "<p>drop-distortion\n<a href=\"https://arxiv.org/pdf/1702.07508.pdf\">https://arxiv.org/pdf/1702.07508.pdf</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fdbe0c609d2f6e881203780e458191c35%2FSelection_068.png?generation=1580484634022137&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 724568,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-21T08:59:39.887000",
      "content": "<p>artificial handwriting\n<a href=\"https://distill.pub/2016/handwriting/\">https://distill.pub/2016/handwriting/</a>\n<a href=\"http://otoro.net/ml/\">http://otoro.net/ml/</a>\n<img src=\"http://blog.otoro.net/wp-content/uploads/sites/2/2015/12/cover_simple_full.svg\" alt=\"\"></p>\n\n<p><a href=\"https://genekogan.com/works/a-book-from-the-sky/\">https://genekogan.com/works/a-book-from-the-sky/</a></p>\n\n<p><img src=\"https://genekogan.com/images/a-book-from-the-sky/radinterpolations/rad30is.gif\" alt=\"\">\n<img src=\"https://cdn.rawgit.com/hardmaru/resnet-cppn-gan-tensorflow/master/examples/example_sinusoid.gif\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 724419,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-21T06:17:17.547000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 722316,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-18T11:44:21.563000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 722322,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-18T11:50:32.477000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 724852,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-21T14:55:07.983000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 721139,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-17T05:04:47.173000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 706961,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-31T04:48:04.067000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 708134,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-02T02:52:50.247000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 708607,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-02T13:25:20.540000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 760537,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-01T12:41:06.157000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 739298,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-07T16:58:18.873000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 730971,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-28T07:27:19.413000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 729042,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-25T16:24:47.597000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 724604,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-21T09:37:23.783000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 724607,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-21T09:47:44.963000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 719277,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-15T10:35:04.510000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 719306,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-15T11:14:16.930000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 707041,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-31T07:53:57.780000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 707078,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-31T08:29:54.473000",
          "content": "",
          "votes": 10,
          "replies": []
        },
        {
          "id": 707248,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-31T13:29:34.917000",
          "content": "",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 706324,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-30T08:40:58.423000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 717410,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-13T04:46:23.400000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 718243,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-14T08:04:35.903000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 706924,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-31T02:31:43.057000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 775962,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-17T03:43:12.780000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 770235,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-12T18:04:17.307000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 770233,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-12T18:02:41.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 770227,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-12T17:51:34.303000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 737229,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-05T04:23:30.307000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 737219,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-05T04:06:46.347000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 717619,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-13T11:11:14.553000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 718375,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-14T10:38:05.637000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 718381,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-14T10:40:48.763000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 719045,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-15T04:07:56.087000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 717306,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-13T02:15:26.040000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 714390,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-09T11:48:55.480000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 706664,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-30T17:15:39.820000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 761708,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-02T21:19:23.920000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 759337,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-28T22:49:45.423000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 722257,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-18T10:21:43.973000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 706310,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-30T08:30:20.873000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "706279": "version 20191230:\n- densenet 121 ensemble + swa without cyclic rate, without bn refinement\n- inference at 25 min\n\n---\n\nversion 20200111:\n- se-resnext50  + balanced sampler + mixup (please refer to readme.ppt at the google drive\n \n\n\n---\n\nversion 20200224:\n- modified se-resnext50 with augmentation that works. please see read file in folder. This version can gives CV 0.991 and LB 0.980  \n \n\n---\n\n\n\nhttps://drive.google.com/open?id=1A5CygainZ4rO_rOs4mUjQk378WlF8MrN\n\n... to be updated ...\ne.g. \n- next version mixup,cutout, manifold mixup, adversarial loss?\nsee https://github.com/rois-codh/kmnist\n       \n- metric learning loss ... large-margin softmax, cosine, center loss?\n\n- long tail, distribution, loss weighing, balance sampling, etc?\n\n- sub class, metric distance without triplet\n\n- attention-based (stroke based?), human in-the-loop\n\n- LSTM, RNN, CTC loss, stroke parsing?\n\n- spatial transformation net, bounding box normalisation \n- augmentation, add random  distractor strokes\n- treating as segmentation problem, or pixel classification + novel pooling\n- feature embedding (https://github.com/qychen13/DifficultyAwareEmbedding), metric learning (https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109)\n\n- fine grained, attribute learning\n- multi-task and task dependency",
    "771046": "this is one of the kaggle competitions that i did the most number of experiments. I notice something interesting.\n\nfor several experiments, it is important to apply the method at the early start of training. e.g. you get different results if you apply ohem on trained models (i.e. end stage of training) versus you apply it right at the start of training.\n\ni think that when data size is small, the number of good solutions (or generalized solutions) region is small. i.e. there are many local minimum. It is difficult to jump from one solution to another solution \n",
    "756402": "this is a reference  version for those who are struggling:\n\nversion 20200224:\n- modified se-resnext50 (64x112 small input) with augmentation that works. please see readme file in folder. \n  This version can gives CV 0.991 and LB 0.980 after you optimized the augmentation hyper parameters  yourselves.\n \n- it provides reference log file for you to compare the loss curve\n\n- it has train/validation split that has gap of CV/LB 0.011 \n\n- code is not complete, but should have enough details to reproduce the above mentioned results.\n  (do not request me for missing files or functions)",
    "737223": "i sometime have discussions with friends on their phd research work. recently, they give me a few pointers on network design: \n\n1. it is possible that a single batch norm cannot normalized the feature well for multi-task. different task needs different normalization.  Care needs to be taken at designing the head\n\n2. you can force the network to learn high frequency features (e.g.  edge, etc) over texture  by adding loss related to high frequency feature extraction , together with the usual cross entropy classification loss.",
    "723886": "magic net\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd43e2775c784672d47d6b37cc582cb90%2FSelection_073.png?generation=1579533369553774&amp;alt=media)\n\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fd5ff591bb2d20c7c87753e6c2170189b%2FSelection_079.png?generation=1579594728840307&amp;alt=media)\n",
    "707588": "after a series of experiments, how are my conclusion:\n\n1. final first LB ranking: 0.99+ after 3 months\n2. shakeup: +/- 0.002 (should be quite small)\n\nbaseline results:\n a. 0.9675~0.9715\n  - standard method (joint 3 class) and ensemble standard model (e.g. densenet, resnet, efficientnet, etc) \n  - to maximize performance, search for best input size (e.g. original size, enlarged size like 224x224,256x256, or even larger), try better augmentation, train longer with better learning rate (swa, cyclic annealing, ...), etc\n  - size, scale normalization\n \n b. 0.9700~0.9800\n... to be updated ... (better formulation, loss etc?)\n\n c. 0.9800~0.9900\n... to be updated ... (some way to create more data?)\n\ndon't forget  Google Cloud AutoML Vision as part of solution !!!!",
    "758757": "good for augmentation?\n\n![](https://raw.githubusercontent.com/warbean/tps_stn_pytorch/master/demo/top_1.gif)\n\nhttps://github.com/WarBean/tps_stn_pytorch",
    "724848": "in google doddle, you can add and subtract drawing. I wonder if this can be done in bengali grapheme? e.g. add/remove or replace vowel, etc\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F249a956c438d0dd9ef6f618882dde10a%2FSelection_087.png?generation=1579618261743829&amp;alt=media)\n\n\n",
    "723868": "new augmentation\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fb35fa0d1187ed81acb31b38fedbc6828%2FSelection_069.png?generation=1579532069051469&amp;alt=media)\n",
    "775716": "private lb of coded posted here:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fda66de721f1d8955b9925e3642f5a0f0%2FSelection_074.png?generation=1584404626455936&amp;alt=media)\n",
    "765573": "there is one extra supervision signal.\n\n```\n list('ক্ট্রো')\n['ক', '্', 'ট', '্', 'র', 'ো']\n```\na list command breaks the graheme into more component. you can use this the measure the amount of common parts (distance) between 2 graphemes ",
    "761546": "if this model is not correlated to image model, it is good for ensemble.\nthis work enables input of very large image\n\nLearning in the Frequency Domain\nKai Xu, Minghai Qin, Fei Sun, Yuhao Wang, Yen-kuang Chen, Fengbo Ren\n\n![](https://storage.googleapis.com/groundai-web-prod/media/users/user_12457/project_409932/images/x2.png)\n\n![](https://storage.googleapis.com/groundai-web-prod/media/users/user_12457/project_409932/images/x4.png)\n\n\n```\n\nResNet-50\t#Channels\tSize Per Channel\tTop-1\tTop-5\tNormalized Input Size\t\nRGB\t            3\t            224x224\t                  75.780\t92.650\t1.0\t \t \nDCT-24 (ours)    24\t            56x56\t                  77.196\t93.504\t0.5\t \t \n \n\n\n```",
    "728658": "top-1 and top-2 recall:\n\n```\nlocal validation results\n\navgerage recall (top-1): 0.982253\n               grapheme_root  0.974616\n             vowel_diacritic  0.992676\n         consonant_diacritic  0.987103\n                    \n\navgerage recall (top-2) : 0.995918\n               grapheme_root  0.993451\n             vowel_diacritic  0.998653\n         consonant_diacritic  0.998116\n                   \n\n```\n\nif you make a mistake, the correct results is probably the top-2. if you can think of a good way to post-process, you can improve your score",
    "757812": "New paper today\nOn Feature Normalization and Data Augmentation\n\nhttps://arxiv.org/pdf/2002.11102.pdf",
    "706574": "What's the meaning of “without bn refinement”? ",
    "706492": "What's the meaning of swa ?",
    "728668": "multi-task and task dependency\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fefb21bf6c5f708e84efb92a2a610f24f%2FClipboard05.png?generation=1579924874232793&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0b65fbdbc95cf5730ca99eff92e7bf6f%2FClipboard04.png?generation=1579924877627631&amp;alt=media)\n",
    "726161": "loss formulation:\n\nsome task (root, constant, or vowel) are more confidence than others.\nsimilarly, some classes are easier within a class.\n\ne.g. given task 1 has class A,B,C,D,E and task 2 has class a,b,c and if we know that some class cannot co-exist, it may give better results if we apply the rules as follow:\n```\n1. naive method:\ntreat task 1 as 5 classes  problem and task 2 as 3 classes. ignore any co-occurrence prior knowledge\n\n 2. use prior knowledge\nif input is a, then it can only be A,B,C\nif input is b, then it can only be A,B,E\nif input is c, then it can only be D,E\n\n```\nif you know the language, you may apply such rules\n",
    "716656": "version 20200111:\nse-resnext50 + balanced sampler + mixup (please refer to readme.ppt at the google drive)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F97673c14b01d4aff29fc1d688b036267%2FSelection_069.png?generation=1578830084131882&amp;alt=media)",
    "706712": "there are 168 grapheme\\_root, 11 vowel\\_diacritic, 7 consonant\\_diacritic and 1295 grapheme classes!\nyou can build a fourth classifier for 1295 grapheme classes and decode the component from it\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbeafecfbcd56e0f8c5ee4ed6d98bd2b3%2FSelection_088.png?generation=1577729205397481&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0e6843ab6fef11c4b7b478c45d0f07f2%2FSelection_087.png?generation=1577729209517703&amp;alt=media)\n",
    "706395": "augmentation in the release 'version 20191230'\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F13e81e022c9b8ebf067517e15d62e938%2FSelection_081.png?generation=1577701870877750&amp;alt=media)\n",
    "716020": "I published a [kernel](https://www.kaggle.com/bibek777/heng-s-starter-training-kernel) based on this so that people can train it using kaggle GPUs",
    "760615": "yet another paper:\nMaxUp: A Simple Way to Improve Generalization of Neural Network Training\n```\n\nMethod\tTop-1 error\tTop-5 error\nVanilla (He et al., 2016a)\t76.3\t-\nDropout (Srivastava et al., 2014)\t76.8\t93.4\nDropPath (Larsson et al., 2017)\t77.1\t93.5\nManifold Mixup (Verma et al., 2019)\t77.5\t93.8\nAutoAugment (Cubuk et al., 2019a)\t77.6\t93.8\nMixup (Zhang et al., 2018)\t77.9\t93.9\nDropBlock (Ghiasi et al., 2018)\t78.3\t94.1\nCutMix (Yun et al., 2019)\t78.6\t94.0\nMaxUp+CutMix\t78.9\t94.2\n\nTable 1: Summary of top1 and top5 accuracies on the validation set of ImageNet for ResNet-50.\n```",
    "758761": "it is possible to treat it as a detection problem\n![](https://raw.githubusercontent.com/RubanSeven/CRAFT_keras/master/images/CRAFT%E9%AB%98%E6%96%AF%E7%83%AD%E5%9B%BE.png)\n\n![](https://pythonawesome.com/content/images/2019/06/CRAFT.jpg)",
    "758090": "https://github.com/sujatasaini/Kuzushiji-DropBlock\n\n \n\n|Models                           | MNIST | Fashion-MNIST | Kuzushiji-MNIST | Kuzushiji-49 |\n|---------------------------------|-------|---------------|-----------------|--------------|\n|[DCNN-DropBlock](DropBlock/Kuzushiji-MNIST/train.py)     | **99.47%** | **93.40%** | **97.66%** | **95.67%** |\n|[DCNN-Dropout](Dropout/train.py)                         | 97.99% | 85.47% | 86.43% | 95.34% |\n|[DCNN-Spatial-Dropout](SpatialDropout/train.py)          | 97.17% | 84.44% |  81.08% | 58.18 |\n ",
    "751629": "another augmentation that works for me:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F5800610b7d3b4f3f03d8e9675b889030%2FSelection_085.png?generation=1582196108850096&amp;alt=media)\n\n![](http://yann.lecun.com/ex/images/invar.png)",
    "737055": "This discussion forum is a gold-mine of ideas. Thanks a lot @hengck23 for initiating this and for sharing such wonderful ideas. \n\nLots to do, lots of learn, not enough time.",
    "733502": "don't forget that each of has has addition 36 unlabelled training images at the test parquet files",
    "733430": "some tricks from ICDAR compeitions\n\nhttps://gateway.newton.ac.uk/sites/default/files/asset/doc/1811/20181127-Large%20Scale%20and%20Unconstrained%20OHCCR%20Based%20on%20Deep%20Learning%20and%20Path%20Signature.pdf",
    "724568": "artificial handwriting\nhttps://distill.pub/2016/handwriting/\nhttp://otoro.net/ml/\n![](http://blog.otoro.net/wp-content/uploads/sites/2/2015/12/cover_simple_full.svg)\n\nhttps://genekogan.com/works/a-book-from-the-sky/\n\n ![](https://genekogan.com/images/a-book-from-the-sky/radinterpolations/rad30is.gif)\n![](https://cdn.rawgit.com/hardmaru/resnet-cppn-gan-tensorflow/master/examples/example_sinusoid.gif)",
    "724419": "https://github.com/osmr/imgclsmob\nmega list of pretrain models",
    "722316": "something interesting ...\nhttps://github.com/Natsu6767/Generating-Devanagari-Using-DRAW\n![](https://github.com/Natsu6767/Generating-Devanagari-Using-DRAW/raw/master/images/devanagari_generate.gif)",
    "721139": "one weird idea to fight overfitting:\n```\n1. given  N class train images\n2. flip the images to make 2 N class (if can do transpose, horizontal, vertical flip to give more class)\n3. pretrain a deep network with 2N class\n\neither \n 4. finetune to original N class\n\nor\n4. continue to 2N class and apply test-time augment (TTA) for test + flip images \n\n```\n\nsince the problem is now more complex, maybe the network is less likely to overfit",
    "706961": "In your code you load the dataset based on `split=balance2/train_b_fold0_184855.npy`. Could you clarify on this? Does it mean that the split fold contains enough samples from each class?",
    "760537": "CamDrop: A New Explanation of Dropout and A Guided Regularization Method for Deep Neural Networks\n\n![](https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/9f13cd4841fb050b345f3a56398871232ea5a58c/3-Figure1-1.png)\n![](https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/9f13cd4841fb050b345f3a56398871232ea5a58c/7-Table6-1.png)",
    "739298": "imageBERT\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F679ae3e3b92efdb848b8a7ef3c4bc81c%2FSelection_067.png?generation=1581094696419061&amp;alt=media)\n",
    "730971": "to my surprise, the error are very subtle\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F910127fdfeb8365828a06159457427df%2FSelection_057.png?generation=1580196437091996&amp;alt=media)\n",
    "729042": "error analysis\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fcc5686df09164b9bded1f56fa081a095%2FSelection_124.png?generation=1579969484208260&amp;alt=media)\n",
    "724604": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F099b1bde79742f98fa45cba0c951d807%2FSelection_085.png?generation=1579599439688221&amp;alt=media)\n",
    "719277": "just an idea\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fbcaa1cc21e04038d9e9dafec16ae41e8%2FSelection_052.png?generation=1579084501734131&amp;alt=media)\n",
    "707041": "Extremely happy to see you in the competition . Considering ,I have joined this competition to try out some basics and some new stuff . Your sharing will surely uplift the purpose of this competition and my own . One question though LSTM would normally help if you have a word written to get something contextual . The few paper I read it is used for that purpose . There could be some that I missed . Here the context that it can derive is if some grapheme combination does not occur in real life and if we can fine-tune the final result based on that ? ",
    "706324": "How much did the swa help? any idea?",
    "706924": "Frog brother, See you again! Thx for your generous share; hope you get better rank",
    "775962": "Thanks for sharing the ideas! Helped a lot and inspired during the competition ",
    "770235": "getting the handwriting strokes\n\nhttps://www.youtube.com/watch?v=47CXR42_b2Y",
    "770233": "just wonder if anyone uses Tesseract?",
    "770227": "this is a common trick to train image of larger size (or deep network) with limited GPU resource:\n\n1) how to train large size:\n1. train image e.g. 360x360\n2. design fully convolutional net\n3. in training crop to smaller size, e.g. 224x224\n4. in final stage, fintunning at full image size (i.e. batch get smaller or freeze bottom layers)\n\n\n2) how to train very deep network\nthis is how vgg16 and googlenett did it when the gpu are small in the early days\n- train few layers first. then freeze these layers. add new trainable layers.\n  (or, implement deep network but alternate between freezing upper and bottom layers at training)\n\n\n",
    "737229": "encoder based augmentation?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0cc99af813bda0e3f1ac4680dd8618ec%2FClipboard09.png?generation=1580876607555657&amp;alt=media)\n",
    "737219": "fine-grained classification\nhttps://arxiv.org/pdf/1903.06150v2.pdf\n\nLooking for the Devil in the Details: Learning Trilinear Attention Sampling\nNetwork for Fine-grained Image Recognition\n\n![](https://paperswithcode.com/media/thumbnails/task/task-0000000719-ccaa52fd.jpg)\n",
    "717619": "Thanks for sharing this.\n\nYou use a file \"grapheme_1295.csv\" for creating the balanced sampler. I have not seen any indication of what the content is or how the content is created. Could you please give some explanations?",
    "717306": "@hengck23 Thanks buddy and best of luck!!!",
    "714390": "Thanks for sharing, I will try to follow your step in this competition. Learned a lot from you ",
    "706664": "I will stay tuned on your post.",
    "761708": "",
    "759337": "",
    "722257": "",
    "706310": ""
  }
}