{
  "id": 101489,
  "title": "[LB 0.4] DenseNet201 with CosFace",
  "url": "/competitions/recursion-cellular-image-classification/discussion/101489",
  "author_name": "",
  "post_date": "2019-07-26T08:35:23.001173900Z",
  "votes": 27,
  "comment_count": 23,
  "views": 0,
  "content": "<h3>[LB 0.4] DenseNet201 with CosFace</h3>\n\n<p>I'm sharing my approach:</p>\n\n<p><strong>Data:</strong>\n* 6-channels images, both sites.\n* Input resolution: 512x512 \n* No external data\n* No TTA</p>\n\n<p><strong>Backbone:</strong> \n* DenseNet 201 (pre-trained on ImageNet) from torchvision package</p>\n\n<p><strong>Loss:</strong> \n* AM-Softmax from CosFace - <a href=\"https://arxiv.org/pdf/1801.09414.pdf\">https://arxiv.org/pdf/1801.09414.pdf</a></p>\n\n<p><strong>Framework:</strong>\n* PyTorch</p>\n\n<p><strong>Local Val Score: 0.59</strong> -&gt; <strong>LB: 0.4</strong></p>",
  "messages": [
    {
      "id": "584626",
      "postDate": "07/26/2019 08:35:23",
      "content": "<h3>[LB 0.4] DenseNet201 with CosFace</h3>\n\n<p>I'm sharing my approach:</p>\n\n<p><strong>Data:</strong>\n* 6-channels images, both sites.\n* Input resolution: 512x512 \n* No external data\n* No TTA</p>\n\n<p><strong>Backbone:</strong> \n* DenseNet 201 (pre-trained on ImageNet) from torchvision package</p>\n\n<p><strong>Loss:</strong> \n* AM-Softmax from CosFace - <a href=\"https://arxiv.org/pdf/1801.09414.pdf\">https://arxiv.org/pdf/1801.09414.pdf</a></p>\n\n<p><strong>Framework:</strong>\n* PyTorch</p>\n\n<p><strong>Local Val Score: 0.59</strong> -&gt; <strong>LB: 0.4</strong></p>",
      "rawMarkdown": "### [LB 0.4] DenseNet201 with CosFace\nI'm sharing my approach:\n\n**Data:**\n* 6-channels images, both sites.\n* Input resolution: 512x512 \n* No external data\n* No TTA\n\n**Backbone:** \n* DenseNet 201 (pre-trained on ImageNet) from torchvision package\n\n**Loss:** \n* AM-Softmax from CosFace - https://arxiv.org/pdf/1801.09414.pdf\n\n**Framework:**\n* PyTorch\n\n**Local Val Score: 0.59** -&gt; **LB: 0.4**",
      "votes": null
    },
    {
      "id": "584683",
      "postDate": "07/26/2019 09:52:44",
      "content": "<p>HI! How many epochs?</p>",
      "rawMarkdown": "HI! How many epochs?",
      "votes": null
    },
    {
      "id": "584692",
      "postDate": "07/26/2019 10:14:21",
      "content": "<p>25+ epochs.</p>",
      "rawMarkdown": "25+ epochs.",
      "votes": null
    },
    {
      "id": "584706",
      "postDate": "07/26/2019 10:55:50",
      "content": "<p>Hello, thanks for sharing your approach!\nHave you tried training the same model with vanilla softmax loss at first? How much improvement does the AM-Softmax loss make?</p>",
      "rawMarkdown": "Hello, thanks for sharing your approach!\nHave you tried training the same model with vanilla softmax loss at first? How much improvement does the AM-Softmax loss make?",
      "votes": null
    },
    {
      "id": "584764",
      "postDate": "07/26/2019 12:47:20",
      "content": "<p>With vanilla Softmax, the same model 0.36. With AM -&gt; 0.4. I am sure that hyperparameters are still not optimal for AM, that's why improvement can be much larger.</p>",
      "rawMarkdown": "With vanilla Softmax, the same model 0.36. With AM -&gt; 0.4. I am sure that hyperparameters are still not optimal for AM, that's why improvement can be much larger.",
      "votes": null
    },
    {
      "id": "584799",
      "postDate": "07/26/2019 13:56:32",
      "content": "<p>hi, thanks for sharing! how are you splitting your validation set? i took the bottom N experiments from train where N = half the number of test experiments based on cells </p>",
      "rawMarkdown": "hi, thanks for sharing! how are you splitting your validation set? i took the bottom N experiments from train where N = half the number of test experiments based on cells",
      "votes": null
    },
    {
      "id": "584821",
      "postDate": "07/26/2019 14:20:28",
      "content": "<p>My split is done by <code>train_test_split()</code> function from <em>scikit-learn</em> with stratification by sirna labels. 95 % of data is in train, 5 % in val. Validation set is not very representative, but I found that even with this amount of training data DenseNet201 overfits, that's why I can't use more data for validation (which could possibly estimate LB scores more precisely).</p>",
      "rawMarkdown": "My split is done by `train_test_split()` function from *scikit-learn* with stratification by sirna labels. 95 % of data is in train, 5 % in val. Validation set is not very representative, but I found that even with this amount of training data DenseNet201 overfits, that's why I can't use more data for validation (which could possibly estimate LB scores more precisely).",
      "votes": null
    },
    {
      "id": "584830",
      "postDate": "07/26/2019 14:30:19",
      "content": "<p>i think you meant 5%/0.05. why not try increase the validation size? Im doing the following for a particular cell type <code>HEPG2</code>.\n<code>\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n</code></p>",
      "rawMarkdown": "i think you meant 5%/0.05. why not try increase the validation size? Im doing the following for a particular cell type `HEPG2`.\n```\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n```",
      "votes": null
    },
    {
      "id": "584836",
      "postDate": "07/26/2019 14:39:19",
      "content": "<p>Thank you. Correct = 5 % in val. Technically I can increase it, but, logically - more data in training -&gt; better generalization on new data.</p>",
      "rawMarkdown": "Thank you. Correct = 5 % in val. Technically I can increase it, but, logically - more data in training -&gt; better generalization on new data.",
      "votes": null
    },
    {
      "id": "584943",
      "postDate": "07/26/2019 17:47:34",
      "content": "<p>Thanks for sharing!</p>\n\n<blockquote>\n  <p>AM-Softmax from CosFace - <a href=\"https://arxiv.org/pdf/1801.09414.pdf\">https://arxiv.org/pdf/1801.09414.pdf</a></p>\n</blockquote>\n\n<p>Can't find \"AM-Softmax\" in this paper, did you mean A-Sotfmax, also studied in this paper and introduced in <a href=\"https://arxiv.org/abs/1704.08063\">https://arxiv.org/abs/1704.08063</a> ?</p>",
      "rawMarkdown": "Thanks for sharing!\n\n&gt; AM-Softmax from CosFace - https://arxiv.org/pdf/1801.09414.pdf\n\nCan't find \"AM-Softmax\" in this paper, did you mean A-Sotfmax, also studied in this paper and introduced in https://arxiv.org/abs/1704.08063 ?",
      "votes": null
    },
    {
      "id": "584959",
      "postDate": "07/26/2019 18:16:54",
      "content": "<p><a href=\"/lopuhin\">@lopuhin</a> sorry for confusing, maybe I wrongly call it AM-softmax. Basically, I used loss from CosFace paper. You also mentioned A-Softmax, so let me clarify a little bit.</p>\n\n<p>The main difference between ArcFace, CosFace and SphereFace (A-Softmax) is a margin type. So for SphereFace (A-Softmax) they used multiplicative margin, and for the ArcFace &amp; CosFace the authors used additive type of margin.</p>\n\n<p>As far as I know from Face Recognition field, additive margin type is the best one, cause it's much easier to achieve network convergence and better separability for embeddings.</p>",
      "rawMarkdown": "lopuhin sorry for confusing, maybe I wrongly call it AM-softmax. Basically, I used loss from CosFace paper. You also mentioned A-Softmax, so let me clarify a little bit.\n\nThe main difference between ArcFace, CosFace and SphereFace (A-Softmax) is a margin type. So for SphereFace (A-Softmax) they used multiplicative margin, and for the ArcFace &amp; CosFace the authors used additive type of margin.\n\nAs far as I know from Face Recognition field, additive margin type is the best one, cause it's much easier to achieve network convergence and better separability for embeddings.",
      "votes": null
    },
    {
      "id": "585194",
      "postDate": "07/27/2019 05:20:09",
      "content": "<p>AM-Softmax and A-Softmax are different works, and the former can be seen as an improvement to the latter.</p>",
      "rawMarkdown": "AM-Softmax and A-Softmax are different works, and the former can be seen as an improvement to the latter.",
      "votes": null
    },
    {
      "id": "585279",
      "postDate": "07/27/2019 07:43:35",
      "content": "<p><a href=\"/alexgruzdev\">@alexgruzdev</a> I am using a DenseNet169 with AM-Softmax from CosFace loss with a lr=3e-4, and epochs=10 on Kaggle Kernels, and the convergence seems to be very slow, is it the same with you?\nThe current score of mine on LB is based on DenseNet169 with CELoss, and same configs. I replace the 1st layer of DenseNet to accomodate 6 channels. \nAny suggestions on how can I improve my score?</p>",
      "rawMarkdown": "alexgruzdev I am using a DenseNet169 with AM-Softmax from CosFace loss with a lr=3e-4, and epochs=10 on Kaggle Kernels, and the convergence seems to be very slow, is it the same with you?\nThe current score of mine on LB is based on DenseNet169 with CELoss, and same configs. I replace the 1st layer of DenseNet to accomodate 6 channels. \nAny suggestions on how can I improve my score?",
      "votes": null
    },
    {
      "id": "585490",
      "postDate": "07/27/2019 14:45:43",
      "content": "<p><a href=\"/ashishsinhaiitr\">@ashishsinhaiitr</a> My current understanding that for any network with CE or AM-Softmax for 20-30 epochs you should achieve at least 0.3 LB score if you properly deal with data. So most probably, you problem is data handling. And you should definitely try more epochs for training (10 is not enough).</p>",
      "rawMarkdown": "ashishsinhaiitr My current understanding that for any network with CE or AM-Softmax for 20-30 epochs you should achieve at least 0.3 LB score if you properly deal with data. So most probably, you problem is data handling. And you should definitely try more epochs for training (10 is not enough).",
      "votes": null
    },
    {
      "id": "585498",
      "postDate": "07/27/2019 14:56:37",
      "content": "<p>i think we both share the same problem of kaggle kernel limit of 9 hours. I am also using <code>densenet121</code> for cell-type specific classifier and still could not achieve 0.3 LB. Does data handling mean use both sites or making use of controls?</p>",
      "rawMarkdown": "i think we both share the same problem of kaggle kernel limit of 9 hours. I am also using `densenet121` for cell-type specific classifier and still could not achieve 0.3 LB. Does data handling mean use both sites or making use of controls?",
      "votes": null
    },
    {
      "id": "585526",
      "postDate": "07/27/2019 15:30:38",
      "content": "<p>If your problem is kaggle kernel limit - you can decrease 512x512 resolution for example to 256x256. I didn't try that, but in <em>ArcFace - Help needed</em> <a href=\"/marekwyborski\">@marekwyborski</a> wrote that for him even with low resolution images it's working OK. With low resolution it should train faster. As for data handling you need to do augmentations and, of course, use both sites.</p>",
      "rawMarkdown": "If your problem is kaggle kernel limit - you can decrease 512x512 resolution for example to 256x256. I didn't try that, but in *ArcFace - Help needed* @marekwyborski wrote that for him even with low resolution images it's working OK. With low resolution it should train faster. As for data handling you need to do augmentations and, of course, use both sites.",
      "votes": null
    },
    {
      "id": "585537",
      "postDate": "07/27/2019 15:50:54",
      "content": "<p>thanks you grib0ed0v. Sorry for the noob question, I thought augmentation is a form of regularization, is it not? so that's normal to help prevent overfitting. i will still try it out. thank you :) </p>",
      "rawMarkdown": "thanks you grib0ed0v. Sorry for the noob question, I thought augmentation is a form of regularization, is it not? so that's normal to help prevent overfitting. i will still try it out. thank you :)",
      "votes": null
    },
    {
      "id": "585976",
      "postDate": "07/28/2019 12:32:29",
      "content": "<p>I am now training for around 25 epochs, and using the dataset prepared by <a href=\"/xhlulu\">@xhlulu</a> which has 224px and in jpeg format.</p>",
      "rawMarkdown": "I am now training for around 25 epochs, and using the dataset prepared by @xhlulu which has 224px and in jpeg format.",
      "votes": null
    },
    {
      "id": "590820",
      "postDate": "08/02/2019 16:47:10",
      "content": "<p>Hi, may I ask? How do you normalize data?</p>",
      "rawMarkdown": "Hi, may I ask? How do you normalize data?",
      "votes": null
    },
    {
      "id": "591073",
      "postDate": "08/03/2019 05:08:33",
      "content": "<p><a href=\"/alexgruzdev\">@alexgruzdev</a> When you use <code>train_test_split</code>, are you using both sites also for validation, or are you keeping one site for validation and one site for training? Thanks</p>",
      "rawMarkdown": "alexgruzdev When you use `train_test_split`, are you using both sites also for validation, or are you keeping one site for validation and one site for training? Thanks",
      "votes": null
    },
    {
      "id": "592363",
      "postDate": "08/05/2019 07:43:23",
      "content": "<p><a href=\"/lorenzofabbri92\">@lorenzofabbri92</a> I am using both sites for validation as well, since if you keep one site for validation and the other one for training - this is famous <strong>\"make the test train again\"</strong> example :)</p>",
      "rawMarkdown": "lorenzofabbri92 I am using both sites for validation as well, since if you keep one site for validation and the other one for training - this is famous **\"make the test train again\"** example :)",
      "votes": null
    },
    {
      "id": "592367",
      "postDate": "08/05/2019 07:49:03",
      "content": "<p><code>T.ToTensor()</code> for transforms + <code>nn.BatchNorm2d(6)</code> as a first layer of my network</p>",
      "rawMarkdown": "`T.ToTensor()` for transforms + `nn.BatchNorm2d(6)` as a first layer of my network",
      "votes": null
    },
    {
      "id": "594781",
      "postDate": "08/08/2019 12:40:57",
      "content": "<p>Hello! May I ask how much time does it take to train 1 epoch of your densenet on 512x512 images? It takes me about 22 minutes on 4 very good gpus -_-\nAnd when I train on 350x350 images it's about 13 minutes.</p>",
      "rawMarkdown": "Hello! May I ask how much time does it take to train 1 epoch of your densenet on 512x512 images? It takes me about 22 minutes on 4 very good gpus -_-\nAnd when I train on 350x350 images it's about 13 minutes.",
      "votes": null
    },
    {
      "id": "595483",
      "postDate": "08/09/2019 09:17:22",
      "content": "<p><a href=\"/rafailfridman\">@rafailfridman</a> I changed resolution for a <strong>448x448</strong>, and basically I have comparable numbers but with only 2 GPUs ~ <strong>20 minutes.</strong> Double-check what <code>nvidia-smi</code> shows for your case. Maybe your bottleneck in data preparation/augmentation and that's why you don't utilize full GPU computations.</p>",
      "rawMarkdown": "rafailfridman I changed resolution for a **448x448**, and basically I have comparable numbers but with only 2 GPUs ~ **20 minutes.** Double-check what `nvidia-smi` shows for your case. Maybe your bottleneck in data preparation/augmentation and that's why you don't utilize full GPU computations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 584683,
      "author_name": "wanizz",
      "author_url": "",
      "post_date": "07/26/2019 09:52:44",
      "content": "<p>HI! How many epochs?</p>",
      "votes": null,
      "replies": [
        {
          "id": 584692,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "07/26/2019 10:14:21",
          "content": "<p>25+ epochs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 584706,
      "author_name": "udaykamal",
      "author_url": "",
      "post_date": "07/26/2019 10:55:50",
      "content": "<p>Hello, thanks for sharing your approach!\nHave you tried training the same model with vanilla softmax loss at first? How much improvement does the AM-Softmax loss make?</p>",
      "votes": null,
      "replies": [
        {
          "id": 584764,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "07/26/2019 12:47:20",
          "content": "<p>With vanilla Softmax, the same model 0.36. With AM -&gt; 0.4. I am sure that hyperparameters are still not optimal for AM, that's why improvement can be much larger.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 584799,
      "author_name": "wjshenggggg",
      "author_url": "",
      "post_date": "07/26/2019 13:56:32",
      "content": "<p>hi, thanks for sharing! how are you splitting your validation set? i took the bottom N experiments from train where N = half the number of test experiments based on cells </p>",
      "votes": null,
      "replies": [
        {
          "id": 584821,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "07/26/2019 14:20:28",
          "content": "<p>My split is done by <code>train_test_split()</code> function from <em>scikit-learn</em> with stratification by sirna labels. 95 % of data is in train, 5 % in val. Validation set is not very representative, but I found that even with this amount of training data DenseNet201 overfits, that's why I can't use more data for validation (which could possibly estimate LB scores more precisely).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584830,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/26/2019 14:30:19",
          "content": "<p>i think you meant 5%/0.05. why not try increase the validation size? Im doing the following for a particular cell type <code>HEPG2</code>.\n<code>\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 584836,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "07/26/2019 14:39:19",
          "content": "<p>Thank you. Correct = 5 % in val. Technically I can increase it, but, logically - more data in training -&gt; better generalization on new data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 584943,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "07/26/2019 17:47:34",
      "content": "<p>Thanks for sharing!</p>\n\n<blockquote>\n  <p>AM-Softmax from CosFace - <a href=\"https://arxiv.org/pdf/1801.09414.pdf\">https://arxiv.org/pdf/1801.09414.pdf</a></p>\n</blockquote>\n\n<p>Can't find \"AM-Softmax\" in this paper, did you mean A-Sotfmax, also studied in this paper and introduced in <a href=\"https://arxiv.org/abs/1704.08063\">https://arxiv.org/abs/1704.08063</a> ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 584959,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "07/26/2019 18:16:54",
          "content": "<p><a href=\"/lopuhin\">@lopuhin</a> sorry for confusing, maybe I wrongly call it AM-softmax. Basically, I used loss from CosFace paper. You also mentioned A-Softmax, so let me clarify a little bit.</p>\n\n<p>The main difference between ArcFace, CosFace and SphereFace (A-Softmax) is a margin type. So for SphereFace (A-Softmax) they used multiplicative margin, and for the ArcFace &amp; CosFace the authors used additive type of margin.</p>\n\n<p>As far as I know from Face Recognition field, additive margin type is the best one, cause it's much easier to achieve network convergence and better separability for embeddings.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585194,
          "author_name": "alexanderliao",
          "author_url": "",
          "post_date": "07/27/2019 05:20:09",
          "content": "<p>AM-Softmax and A-Softmax are different works, and the former can be seen as an improvement to the latter.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 594781,
          "author_name": "rafailfridman",
          "author_url": "",
          "post_date": "08/08/2019 12:40:57",
          "content": "<p>Hello! May I ask how much time does it take to train 1 epoch of your densenet on 512x512 images? It takes me about 22 minutes on 4 very good gpus -_-\nAnd when I train on 350x350 images it's about 13 minutes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 595483,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "08/09/2019 09:17:22",
          "content": "<p><a href=\"/rafailfridman\">@rafailfridman</a> I changed resolution for a <strong>448x448</strong>, and basically I have comparable numbers but with only 2 GPUs ~ <strong>20 minutes.</strong> Double-check what <code>nvidia-smi</code> shows for your case. Maybe your bottleneck in data preparation/augmentation and that's why you don't utilize full GPU computations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 585279,
      "author_name": "ashishsinhaiitr",
      "author_url": "",
      "post_date": "07/27/2019 07:43:35",
      "content": "<p><a href=\"/alexgruzdev\">@alexgruzdev</a> I am using a DenseNet169 with AM-Softmax from CosFace loss with a lr=3e-4, and epochs=10 on Kaggle Kernels, and the convergence seems to be very slow, is it the same with you?\nThe current score of mine on LB is based on DenseNet169 with CELoss, and same configs. I replace the 1st layer of DenseNet to accomodate 6 channels. \nAny suggestions on how can I improve my score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 585490,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "07/27/2019 14:45:43",
          "content": "<p><a href=\"/ashishsinhaiitr\">@ashishsinhaiitr</a> My current understanding that for any network with CE or AM-Softmax for 20-30 epochs you should achieve at least 0.3 LB score if you properly deal with data. So most probably, you problem is data handling. And you should definitely try more epochs for training (10 is not enough).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585498,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/27/2019 14:56:37",
          "content": "<p>i think we both share the same problem of kaggle kernel limit of 9 hours. I am also using <code>densenet121</code> for cell-type specific classifier and still could not achieve 0.3 LB. Does data handling mean use both sites or making use of controls?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585526,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "07/27/2019 15:30:38",
          "content": "<p>If your problem is kaggle kernel limit - you can decrease 512x512 resolution for example to 256x256. I didn't try that, but in <em>ArcFace - Help needed</em> <a href=\"/marekwyborski\">@marekwyborski</a> wrote that for him even with low resolution images it's working OK. With low resolution it should train faster. As for data handling you need to do augmentations and, of course, use both sites.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585537,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/27/2019 15:50:54",
          "content": "<p>thanks you grib0ed0v. Sorry for the noob question, I thought augmentation is a form of regularization, is it not? so that's normal to help prevent overfitting. i will still try it out. thank you :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585976,
          "author_name": "ashishsinhaiitr",
          "author_url": "",
          "post_date": "07/28/2019 12:32:29",
          "content": "<p>I am now training for around 25 epochs, and using the dataset prepared by <a href=\"/xhlulu\">@xhlulu</a> which has 224px and in jpeg format.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 591073,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/03/2019 05:08:33",
          "content": "<p><a href=\"/alexgruzdev\">@alexgruzdev</a> When you use <code>train_test_split</code>, are you using both sites also for validation, or are you keeping one site for validation and one site for training? Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592363,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "08/05/2019 07:43:23",
          "content": "<p><a href=\"/lorenzofabbri92\">@lorenzofabbri92</a> I am using both sites for validation as well, since if you keep one site for validation and the other one for training - this is famous <strong>\"make the test train again\"</strong> example :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 590820,
      "author_name": "joven1997",
      "author_url": "",
      "post_date": "08/02/2019 16:47:10",
      "content": "<p>Hi, may I ask? How do you normalize data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 592367,
          "author_name": "alexgruzdev",
          "author_url": "",
          "post_date": "08/05/2019 07:49:03",
          "content": "<p><code>T.ToTensor()</code> for transforms + <code>nn.BatchNorm2d(6)</code> as a first layer of my network</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "584626": "### [LB 0.4] DenseNet201 with CosFace\nI'm sharing my approach:\n\n**Data:**\n* 6-channels images, both sites.\n* Input resolution: 512x512 \n* No external data\n* No TTA\n\n**Backbone:** \n* DenseNet 201 (pre-trained on ImageNet) from torchvision package\n\n**Loss:** \n* AM-Softmax from CosFace - https://arxiv.org/pdf/1801.09414.pdf\n\n**Framework:**\n* PyTorch\n\n**Local Val Score: 0.59** -&gt; **LB: 0.4**",
    "584683": "HI! How many epochs?",
    "584692": "25+ epochs.",
    "584706": "Hello, thanks for sharing your approach!\nHave you tried training the same model with vanilla softmax loss at first? How much improvement does the AM-Softmax loss make?",
    "584764": "With vanilla Softmax, the same model 0.36. With AM -&gt; 0.4. I am sure that hyperparameters are still not optimal for AM, that's why improvement can be much larger.",
    "584799": "hi, thanks for sharing! how are you splitting your validation set? i took the bottom N experiments from train where N = half the number of test experiments based on cells",
    "584821": "My split is done by `train_test_split()` function from *scikit-learn* with stratification by sirna labels. 95 % of data is in train, 5 % in val. Validation set is not very representative, but I found that even with this amount of training data DenseNet201 overfits, that's why I can't use more data for validation (which could possibly estimate LB scores more precisely).",
    "584830": "i think you meant 5%/0.05. why not try increase the validation size? Im doing the following for a particular cell type `HEPG2`.\n```\ntrain len: 5536      experiments: ['HEPG2-01' 'HEPG2-02' 'HEPG2-03' 'HEPG2-04' 'HEPG2-05']\nvalid len: 2214      experiments: ['HEPG2-06' 'HEPG2-07']\ntest  len: 4429      experiments: ['HEPG2-08' 'HEPG2-09' 'HEPG2-10' 'HEPG2-11']\n```",
    "584836": "Thank you. Correct = 5 % in val. Technically I can increase it, but, logically - more data in training -&gt; better generalization on new data.",
    "584943": "Thanks for sharing!\n\n&gt; AM-Softmax from CosFace - https://arxiv.org/pdf/1801.09414.pdf\n\nCan't find \"AM-Softmax\" in this paper, did you mean A-Sotfmax, also studied in this paper and introduced in https://arxiv.org/abs/1704.08063 ?",
    "584959": "lopuhin sorry for confusing, maybe I wrongly call it AM-softmax. Basically, I used loss from CosFace paper. You also mentioned A-Softmax, so let me clarify a little bit.\n\nThe main difference between ArcFace, CosFace and SphereFace (A-Softmax) is a margin type. So for SphereFace (A-Softmax) they used multiplicative margin, and for the ArcFace &amp; CosFace the authors used additive type of margin.\n\nAs far as I know from Face Recognition field, additive margin type is the best one, cause it's much easier to achieve network convergence and better separability for embeddings.",
    "585194": "AM-Softmax and A-Softmax are different works, and the former can be seen as an improvement to the latter.",
    "585279": "alexgruzdev I am using a DenseNet169 with AM-Softmax from CosFace loss with a lr=3e-4, and epochs=10 on Kaggle Kernels, and the convergence seems to be very slow, is it the same with you?\nThe current score of mine on LB is based on DenseNet169 with CELoss, and same configs. I replace the 1st layer of DenseNet to accomodate 6 channels. \nAny suggestions on how can I improve my score?",
    "585490": "ashishsinhaiitr My current understanding that for any network with CE or AM-Softmax for 20-30 epochs you should achieve at least 0.3 LB score if you properly deal with data. So most probably, you problem is data handling. And you should definitely try more epochs for training (10 is not enough).",
    "585498": "i think we both share the same problem of kaggle kernel limit of 9 hours. I am also using `densenet121` for cell-type specific classifier and still could not achieve 0.3 LB. Does data handling mean use both sites or making use of controls?",
    "585526": "If your problem is kaggle kernel limit - you can decrease 512x512 resolution for example to 256x256. I didn't try that, but in *ArcFace - Help needed* @marekwyborski wrote that for him even with low resolution images it's working OK. With low resolution it should train faster. As for data handling you need to do augmentations and, of course, use both sites.",
    "585537": "thanks you grib0ed0v. Sorry for the noob question, I thought augmentation is a form of regularization, is it not? so that's normal to help prevent overfitting. i will still try it out. thank you :)",
    "585976": "I am now training for around 25 epochs, and using the dataset prepared by @xhlulu which has 224px and in jpeg format.",
    "590820": "Hi, may I ask? How do you normalize data?",
    "591073": "alexgruzdev When you use `train_test_split`, are you using both sites also for validation, or are you keeping one site for validation and one site for training? Thanks",
    "592363": "lorenzofabbri92 I am using both sites for validation as well, since if you keep one site for validation and the other one for training - this is famous **\"make the test train again\"** example :)",
    "592367": "`T.ToTensor()` for transforms + `nn.BatchNorm2d(6)` as a first layer of my network",
    "594781": "Hello! May I ask how much time does it take to train 1 epoch of your densenet on 512x512 images? It takes me about 22 minutes on 4 very good gpus -_-\nAnd when I train on 350x350 images it's about 13 minutes.",
    "595483": "rafailfridman I changed resolution for a **448x448**, and basically I have comparable numbers but with only 2 GPUs ~ **20 minutes.** Double-check what `nvidia-smi` shows for your case. Maybe your bottleneck in data preparation/augmentation and that's why you don't utilize full GPU computations."
  },
  "source": "meta"
}