{
  "id": 172870,
  "title": "Changing the drop_connect_rate argument improves my EfficienNet model ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172870",
  "author_name": "",
  "post_date": "2020-08-06T20:02:27.161963500Z",
  "votes": 22,
  "comment_count": 8,
  "views": 0,
  "content": "<p>By default the <code>drop_connect_rate</code> argument value it's 0.2, i decide to change the value to 0.4 and my CV/LB are higher. Maybe this can be helpful for someone.</p>\n<p><code>model = EfficientNetB0(drop_connect_rate = 0.4, ...)</code></p>\n<p>==================================================<br>\nUPDATE: Example<br>\n**From the excellent example of ** <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<pre><code>build_model(dim = 128, ef = 0):\n    inp = tf.keras.layers.Input(shape = (dim, dim, 3))\n    base = EFNS[ef](\ninput_shape = (dim, dim, 3),\nweights = 'noisy-student',\n include_top = False,\n drop_connect_rate = 0.4\n)\n\n    # Rebuild top\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(1, activation = 'sigmoid')(x)\n\n    # Compile\n    model = tf.keras.Model(inputs = inp, outputs = x)\n    opt = tf.keras.optimizers.Adam(learning_rate = 0.001)\n    loss = tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.05) \n    model.compile(optimizer = opt, loss = loss, metrics = ['AUC'])\n    return model\n</code></pre>\n<p>NOTE: Don't forget that this is a regularization technique for EfficientNet.</p>",
  "messages": [
    {
      "id": "960944",
      "postDate": "08/06/2020 20:02:27",
      "content": "<p>By default the <code>drop_connect_rate</code> argument value it's 0.2, i decide to change the value to 0.4 and my CV/LB are higher. Maybe this can be helpful for someone.</p>\n<p><code>model = EfficientNetB0(drop_connect_rate = 0.4, ...)</code></p>\n<p>==================================================<br>\nUPDATE: Example<br>\n**From the excellent example of ** <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<pre><code>build_model(dim = 128, ef = 0):\n    inp = tf.keras.layers.Input(shape = (dim, dim, 3))\n    base = EFNS[ef](\ninput_shape = (dim, dim, 3),\nweights = 'noisy-student',\n include_top = False,\n drop_connect_rate = 0.4\n)\n\n    # Rebuild top\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(1, activation = 'sigmoid')(x)\n\n    # Compile\n    model = tf.keras.Model(inputs = inp, outputs = x)\n    opt = tf.keras.optimizers.Adam(learning_rate = 0.001)\n    loss = tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.05) \n    model.compile(optimizer = opt, loss = loss, metrics = ['AUC'])\n    return model\n</code></pre>\n<p>NOTE: Don't forget that this is a regularization technique for EfficientNet.</p>",
      "rawMarkdown": "By default the `drop_connect_rate` argument value it's 0.2, i decide to change the value to 0.4 and my CV/LB are higher. Maybe this can be helpful for someone.\n\n`model = EfficientNetB0(drop_connect_rate = 0.4, ...)`\n\n==================================================\nUPDATE: Example\n**From the excellent example of ** @cdeotte \n\n```python\nbuild_model(dim = 128, ef = 0):\n    inp = tf.keras.layers.Input(shape = (dim, dim, 3))\n    base = EFNS[ef](\ninput_shape = (dim, dim, 3),\nweights = 'noisy-student',\n include_top = False,\n drop_connect_rate = 0.4\n)\n    \n    # Rebuild top\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(1, activation = 'sigmoid')(x)\n    \n    # Compile\n    model = tf.keras.Model(inputs = inp, outputs = x)\n    opt = tf.keras.optimizers.Adam(learning_rate = 0.001)\n    loss = tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.05) \n    model.compile(optimizer = opt, loss = loss, metrics = ['AUC'])\n    return model\n```\n\nNOTE: Don't forget that this is a regularization technique for EfficientNet.",
      "votes": null
    },
    {
      "id": "961305",
      "postDate": "08/07/2020 04:59:50",
      "content": "<p>What does this parameter doing?? <a href=\"/hiramcho\">@hiramcho</a> </p>",
      "rawMarkdown": "What does this parameter doing?? @hiramcho",
      "votes": null
    },
    {
      "id": "961764",
      "postDate": "08/07/2020 13:34:23",
      "content": "<p><a href=\"https://www.kaggle.com/msharuk589\" target=\"_blank\">@msharuk589</a> controls the dropout rate responsible for stochastic depth.  <code>drop_connect_rate</code> serves as a toggle for extra regularization in finetuning, but does not affect loaded weights. </p>",
      "rawMarkdown": "msharuk589 controls the dropout rate responsible for stochastic depth.  `drop_connect_rate` serves as a toggle for extra regularization in finetuning, but does not affect loaded weights.",
      "votes": null
    },
    {
      "id": "961854",
      "postDate": "08/07/2020 15:02:19",
      "content": "<p>For those interested in using it with efficientnet-pytorch 0.6.3 (what you get with pip install) here is some code that work for me.  </p>\n<pre><code>from efficientnet_pytorch.utils import (\n    round_filters,\n    round_repeats,\n    drop_connect,\n    get_same_padding_conv2d,\n    get_model_params,\n    efficientnet_params,\n    load_pretrained_weights,\n    Swish,\n    MemoryEfficientSwish,\n)\n\ndef from_pretrained(model_name, advprop=False, \n                    num_classes=1000, in_channels=3, \n                    drop_connect_rate=0.2):\n    model = EfficientNet.from_name(model_name, \n                                   override_params={'num_classes': num_classes,\n                                                    'drop_connect_rate':drop_connect_rate,\n                                                   })\n    load_pretrained_weights(model, \n                            model_name, \n                            load_fc=(num_classes == 1000), \n                            advprop=advprop)\n    if in_channels != 3:\n        Conv2d = get_same_padding_conv2d(image_size = model._global_params.image_size)\n        out_channels = round_filters(32, model._global_params)\n        model._conv_stem = Conv2d(in_channels, out_channels, kernel_size=3, stride=2, bias=False)\n    return model\n\n\nmodel = from_pretrained('efficientnet-b0', drop_connect_rate=0.4) \n</code></pre>",
      "rawMarkdown": "For those interested in using it with efficientnet-pytorch 0.6.3 (what you get with pip install) here is some code that work for me.  \n\n```\nfrom efficientnet_pytorch.utils import (\n    round_filters,\n    round_repeats,\n    drop_connect,\n    get_same_padding_conv2d,\n    get_model_params,\n    efficientnet_params,\n    load_pretrained_weights,\n    Swish,\n    MemoryEfficientSwish,\n)\n\ndef from_pretrained(model_name, advprop=False, \n                    num_classes=1000, in_channels=3, \n                    drop_connect_rate=0.2):\n    model = EfficientNet.from_name(model_name, \n                                   override_params={'num_classes': num_classes,\n                                                    'drop_connect_rate':drop_connect_rate,\n                                                   })\n    load_pretrained_weights(model, \n                            model_name, \n                            load_fc=(num_classes == 1000), \n                            advprop=advprop)\n    if in_channels != 3:\n        Conv2d = get_same_padding_conv2d(image_size = model._global_params.image_size)\n        out_channels = round_filters(32, model._global_params)\n        model._conv_stem = Conv2d(in_channels, out_channels, kernel_size=3, stride=2, bias=False)\n    return model\n\n\nmodel = from_pretrained('efficientnet-b0', drop_connect_rate=0.4) \n```",
      "votes": null
    },
    {
      "id": "961914",
      "postDate": "08/07/2020 15:51:13",
      "content": "<p>Will try thanks for sharing.</p>",
      "rawMarkdown": "Will try thanks for sharing.",
      "votes": null
    },
    {
      "id": "961937",
      "postDate": "08/07/2020 16:23:48",
      "content": "<p>I forgot to specify that the example it's for Tensorflow 2.x, thanks for share.</p>",
      "rawMarkdown": "I forgot to specify that the example it's for Tensorflow 2.x, thanks for share.",
      "votes": null
    },
    {
      "id": "962062",
      "postDate": "08/07/2020 19:03:55",
      "content": "<p>Mmh tried it and it didn't give me better results, i guess it depends if you re using hard augmentations (ex: Coarse Dropout)</p>",
      "rawMarkdown": "Mmh tried it and it didn't give me better results, i guess it depends if you re using hard augmentations (ex: Coarse Dropout)",
      "votes": null
    },
    {
      "id": "962079",
      "postDate": "08/07/2020 19:25:59",
      "content": "<p>Yes, that's why i added a \"NOTE\" in the post beacause before use it, you must understand your model and know if it needs more regularization. </p>",
      "rawMarkdown": "Yes, that's why i added a \"NOTE\" in the post beacause before use it, you must understand your model and know if it needs more regularization.",
      "votes": null
    },
    {
      "id": "966770",
      "postDate": "08/11/2020 16:48:31",
      "content": "<p>In my case, it works. Thanks, <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a>, it's a great share. -)</p>",
      "rawMarkdown": "In my case, it works. Thanks, @hiramcho, it's a great share. -)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 961854,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/07/2020 15:02:19",
      "content": "<p>For those interested in using it with efficientnet-pytorch 0.6.3 (what you get with pip install) here is some code that work for me.  </p>\n<pre><code>from efficientnet_pytorch.utils import (\n    round_filters,\n    round_repeats,\n    drop_connect,\n    get_same_padding_conv2d,\n    get_model_params,\n    efficientnet_params,\n    load_pretrained_weights,\n    Swish,\n    MemoryEfficientSwish,\n)\n\ndef from_pretrained(model_name, advprop=False, \n                    num_classes=1000, in_channels=3, \n                    drop_connect_rate=0.2):\n    model = EfficientNet.from_name(model_name, \n                                   override_params={'num_classes': num_classes,\n                                                    'drop_connect_rate':drop_connect_rate,\n                                                   })\n    load_pretrained_weights(model, \n                            model_name, \n                            load_fc=(num_classes == 1000), \n                            advprop=advprop)\n    if in_channels != 3:\n        Conv2d = get_same_padding_conv2d(image_size = model._global_params.image_size)\n        out_channels = round_filters(32, model._global_params)\n        model._conv_stem = Conv2d(in_channels, out_channels, kernel_size=3, stride=2, bias=False)\n    return model\n\n\nmodel = from_pretrained('efficientnet-b0', drop_connect_rate=0.4) \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 961937,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "08/07/2020 16:23:48",
          "content": "<p>I forgot to specify that the example it's for Tensorflow 2.x, thanks for share.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 961914,
      "author_name": "arroqc",
      "author_url": "",
      "post_date": "08/07/2020 15:51:13",
      "content": "<p>Will try thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 966770,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "08/11/2020 16:48:31",
      "content": "<p>In my case, it works. Thanks, <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a>, it's a great share. -)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 961305,
      "author_name": "msharuk589",
      "author_url": "",
      "post_date": "08/07/2020 04:59:50",
      "content": "<p>What does this parameter doing?? <a href=\"/hiramcho\">@hiramcho</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 961764,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "08/07/2020 13:34:23",
          "content": "<p><a href=\"https://www.kaggle.com/msharuk589\" target=\"_blank\">@msharuk589</a> controls the dropout rate responsible for stochastic depth.  <code>drop_connect_rate</code> serves as a toggle for extra regularization in finetuning, but does not affect loaded weights. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 962062,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "08/07/2020 19:03:55",
      "content": "<p>Mmh tried it and it didn't give me better results, i guess it depends if you re using hard augmentations (ex: Coarse Dropout)</p>",
      "votes": null,
      "replies": [
        {
          "id": 962079,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "08/07/2020 19:25:59",
          "content": "<p>Yes, that's why i added a \"NOTE\" in the post beacause before use it, you must understand your model and know if it needs more regularization. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "960944": "By default the `drop_connect_rate` argument value it's 0.2, i decide to change the value to 0.4 and my CV/LB are higher. Maybe this can be helpful for someone.\n\n`model = EfficientNetB0(drop_connect_rate = 0.4, ...)`\n\n==================================================\nUPDATE: Example\n**From the excellent example of ** @cdeotte \n\n```python\nbuild_model(dim = 128, ef = 0):\n    inp = tf.keras.layers.Input(shape = (dim, dim, 3))\n    base = EFNS[ef](\ninput_shape = (dim, dim, 3),\nweights = 'noisy-student',\n include_top = False,\n drop_connect_rate = 0.4\n)\n    \n    # Rebuild top\n    x = base(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(1, activation = 'sigmoid')(x)\n    \n    # Compile\n    model = tf.keras.Model(inputs = inp, outputs = x)\n    opt = tf.keras.optimizers.Adam(learning_rate = 0.001)\n    loss = tf.keras.losses.BinaryCrossentropy(label_smoothing = 0.05) \n    model.compile(optimizer = opt, loss = loss, metrics = ['AUC'])\n    return model\n```\n\nNOTE: Don't forget that this is a regularization technique for EfficientNet.",
    "961305": "What does this parameter doing?? @hiramcho",
    "961764": "msharuk589 controls the dropout rate responsible for stochastic depth.  `drop_connect_rate` serves as a toggle for extra regularization in finetuning, but does not affect loaded weights.",
    "961854": "For those interested in using it with efficientnet-pytorch 0.6.3 (what you get with pip install) here is some code that work for me.  \n\n```\nfrom efficientnet_pytorch.utils import (\n    round_filters,\n    round_repeats,\n    drop_connect,\n    get_same_padding_conv2d,\n    get_model_params,\n    efficientnet_params,\n    load_pretrained_weights,\n    Swish,\n    MemoryEfficientSwish,\n)\n\ndef from_pretrained(model_name, advprop=False, \n                    num_classes=1000, in_channels=3, \n                    drop_connect_rate=0.2):\n    model = EfficientNet.from_name(model_name, \n                                   override_params={'num_classes': num_classes,\n                                                    'drop_connect_rate':drop_connect_rate,\n                                                   })\n    load_pretrained_weights(model, \n                            model_name, \n                            load_fc=(num_classes == 1000), \n                            advprop=advprop)\n    if in_channels != 3:\n        Conv2d = get_same_padding_conv2d(image_size = model._global_params.image_size)\n        out_channels = round_filters(32, model._global_params)\n        model._conv_stem = Conv2d(in_channels, out_channels, kernel_size=3, stride=2, bias=False)\n    return model\n\n\nmodel = from_pretrained('efficientnet-b0', drop_connect_rate=0.4) \n```",
    "961914": "Will try thanks for sharing.",
    "961937": "I forgot to specify that the example it's for Tensorflow 2.x, thanks for share.",
    "962062": "Mmh tried it and it didn't give me better results, i guess it depends if you re using hard augmentations (ex: Coarse Dropout)",
    "962079": "Yes, that's why i added a \"NOTE\" in the post beacause before use it, you must understand your model and know if it needs more regularization.",
    "966770": "In my case, it works. Thanks, @hiramcho, it's a great share. -)"
  },
  "source": "meta"
}