{
  "id": 315129,
  "title": "some of the modeling of arcface public kernel  are wrong",
  "url": "/competitions/happy-whale-and-dolphin/discussion/315129",
  "author_name": "",
  "post_date": "2022-03-26T11:48:27.711053700Z",
  "votes": 58,
  "comment_count": 8,
  "views": 0,
  "content": "<p>i spend time to investigate why some public kernel are better than others.</p>\n<p>i used one week to discover this:</p>\n<ul>\n<li>norm (batchnorm or llayernorm) before L2 is necessary</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/82409\" target=\"_blank\">https://www.kaggle.com/c/humpback-whale-identification/discussion/82409</a><br>\n\"BatchNorm vs LayerNorm before L2 normalization: LayerNorm better\"</p>\n<p><a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/109987\" target=\"_blank\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/109987</a><br>\n\"Batch normalization layer is vital just before the L2-normalized features multiplication by the weights. It is indeed mentioned in the original article, but I didn't pay attention at first, and ArcFace started working for me only after I added it.\"</p>",
  "messages": [
    {
      "id": "1735555",
      "postDate": "03/26/2022 11:48:27",
      "content": "<p>i spend time to investigate why some public kernel are better than others.</p>\n<p>i used one week to discover this:</p>\n<ul>\n<li>norm (batchnorm or llayernorm) before L2 is necessary</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/82409\" target=\"_blank\">https://www.kaggle.com/c/humpback-whale-identification/discussion/82409</a><br>\n\"BatchNorm vs LayerNorm before L2 normalization: LayerNorm better\"</p>\n<p><a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/109987\" target=\"_blank\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/109987</a><br>\n\"Batch normalization layer is vital just before the L2-normalized features multiplication by the weights. It is indeed mentioned in the original article, but I didn't pay attention at first, and ArcFace started working for me only after I added it.\"</p>",
      "rawMarkdown": "i spend time to investigate why some public kernel are better than others.\n\ni used one week to discover this:\n- norm (batchnorm or llayernorm) before L2 is necessary\n\nhttps://www.kaggle.com/c/humpback-whale-identification/discussion/82409\n\"BatchNorm vs LayerNorm before L2 normalization: LayerNorm better\"\n\nhttps://www.kaggle.com/c/recursion-cellular-image-classification/discussion/109987\n\"Batch normalization layer is vital just before the L2-normalized features multiplication by the weights. It is indeed mentioned in the original article, but I didn't pay attention at first, and ArcFace started working for me only after I added it.\"",
      "votes": null
    },
    {
      "id": "1735584",
      "postDate": "03/26/2022 12:29:26",
      "content": "<p>The characteristics of this game are overfitting the training set. It is useful to add constraints and tricks at the feature level.<br>\nUsing bn layer and feature loss is equivalent to assuming that the feature space, increases the features to distinguish, also eased the overfitting.</p>",
      "rawMarkdown": "The characteristics of this game are overfitting the training set. It is useful to add constraints and tricks at the feature level.\nUsing bn layer and feature loss is equivalent to assuming that the feature space, increases the features to distinguish, also eased the overfitting.",
      "votes": null
    },
    {
      "id": "1735647",
      "postDate": "03/26/2022 13:29:35",
      "content": "<p>i have a feeling that normalisation layer is smiliar is the PCA trick before deep learning metric learning is \"invented\". </p>\n<p>In traditional old ML methods, it is necessary to perform PCA to ensure that the distance is spherical in the embedding space. cosine distance is only meaningful if  the embedding space is spherical</p>\n<p>it is also a known trick that batch size should be large enough for metric learning/retrival/face recognition problem/contrastive learning </p>",
      "rawMarkdown": "i have a feeling that normalisation layer is smiliar is the PCA trick before deep learning metric learning is \"invented\". \n\nIn traditional old ML methods, it is necessary to perform PCA to ensure that the distance is spherical in the embedding space. cosine distance is only meaningful if  the embedding space is spherical\n\nit is also a known trick that batch size should be large enough for metric learning/retrival/face recognition problem/contrastive learning",
      "votes": null
    },
    {
      "id": "1736096",
      "postDate": "03/27/2022 00:28:57",
      "content": "<p>Is this for training or inference or both?  </p>",
      "rawMarkdown": "Is this for training or inference or both?",
      "votes": null
    },
    {
      "id": "1738614",
      "postDate": "03/29/2022 12:04:35",
      "content": "<p>Yep, in <a href=\"https://arxiv.org/pdf/1801.07698.pdf\" target=\"_blank\">original paper</a> they say, that:</p>\n<blockquote>\n  <p>After the last convolutional layer, we explore the BN-Dropout-FC-BN structure to get the final 512-D embedding feature.</p>\n</blockquote>",
      "rawMarkdown": "Yep, in [original paper](https://arxiv.org/pdf/1801.07698.pdf) they say, that:\n> After the last convolutional layer, we explore the BN-Dropout-FC-BN structure to get the final 512-D embedding feature.",
      "votes": null
    },
    {
      "id": "1741804",
      "postDate": "04/01/2022 07:10:07",
      "content": "<p>Batch norm parameters are being updated during training</p>\n<p>In the inference phase, the batch norm parameters called gamma and beta (which are learned in the training phase) and their running mean and std are fixed and will be used for computing the output.</p>",
      "rawMarkdown": "Batch norm parameters are being updated during training\n\nIn the inference phase, the batch norm parameters called gamma and beta (which are learned in the training phase) and their running mean and std are fixed and will be used for computing the output.",
      "votes": null
    },
    {
      "id": "1750734",
      "postDate": "04/10/2022 02:52:51",
      "content": "<p>Thank you for sharing great information!<br>\nCould you show me some example of  kernels better than others, or give me hint to find better notebook.<br>\nI can't judge what notebook is better or not…</p>",
      "rawMarkdown": "Thank you for sharing great information!\nCould you show me some example of  kernels better than others, or give me hint to find better notebook.\nI can't judge what notebook is better or not...",
      "votes": null
    },
    {
      "id": "1750945",
      "postDate": "04/10/2022 08:40:11",
      "content": "<p>Most of the time there is no wrong or right in machine learning. Try all options and see what works best.</p>",
      "rawMarkdown": "Most of the time there is no wrong or right in machine learning. Try all options and see what works best.",
      "votes": null
    },
    {
      "id": "1751128",
      "postDate": "04/10/2022 12:01:00",
      "content": "<p>Indeed. There's still a lot to be studied, so it's hard to say if something is right or wrong.</p>",
      "rawMarkdown": "Indeed. There's still a lot to be studied, so it's hard to say if something is right or wrong.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1735584,
      "author_name": "biglafe",
      "author_url": "",
      "post_date": "03/26/2022 12:29:26",
      "content": "<p>The characteristics of this game are overfitting the training set. It is useful to add constraints and tricks at the feature level.<br>\nUsing bn layer and feature loss is equivalent to assuming that the feature space, increases the features to distinguish, also eased the overfitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1735647,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/26/2022 13:29:35",
          "content": "<p>i have a feeling that normalisation layer is smiliar is the PCA trick before deep learning metric learning is \"invented\". </p>\n<p>In traditional old ML methods, it is necessary to perform PCA to ensure that the distance is spherical in the embedding space. cosine distance is only meaningful if  the embedding space is spherical</p>\n<p>it is also a known trick that batch size should be large enough for metric learning/retrival/face recognition problem/contrastive learning </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1736096,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "03/27/2022 00:28:57",
      "content": "<p>Is this for training or inference or both?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1741804,
          "author_name": "ptran1203",
          "author_url": "",
          "post_date": "04/01/2022 07:10:07",
          "content": "<p>Batch norm parameters are being updated during training</p>\n<p>In the inference phase, the batch norm parameters called gamma and beta (which are learned in the training phase) and their running mean and std are fixed and will be used for computing the output.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1738614,
      "author_name": "vadbeg",
      "author_url": "",
      "post_date": "03/29/2022 12:04:35",
      "content": "<p>Yep, in <a href=\"https://arxiv.org/pdf/1801.07698.pdf\" target=\"_blank\">original paper</a> they say, that:</p>\n<blockquote>\n  <p>After the last convolutional layer, we explore the BN-Dropout-FC-BN structure to get the final 512-D embedding feature.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1750734,
      "author_name": "whitelily",
      "author_url": "",
      "post_date": "04/10/2022 02:52:51",
      "content": "<p>Thank you for sharing great information!<br>\nCould you show me some example of  kernels better than others, or give me hint to find better notebook.<br>\nI can't judge what notebook is better or not…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1750945,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "04/10/2022 08:40:11",
      "content": "<p>Most of the time there is no wrong or right in machine learning. Try all options and see what works best.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1751128,
          "author_name": "robsonsan",
          "author_url": "",
          "post_date": "04/10/2022 12:01:00",
          "content": "<p>Indeed. There's still a lot to be studied, so it's hard to say if something is right or wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1735555": "i spend time to investigate why some public kernel are better than others.\n\ni used one week to discover this:\n- norm (batchnorm or llayernorm) before L2 is necessary\n\nhttps://www.kaggle.com/c/humpback-whale-identification/discussion/82409\n\"BatchNorm vs LayerNorm before L2 normalization: LayerNorm better\"\n\nhttps://www.kaggle.com/c/recursion-cellular-image-classification/discussion/109987\n\"Batch normalization layer is vital just before the L2-normalized features multiplication by the weights. It is indeed mentioned in the original article, but I didn't pay attention at first, and ArcFace started working for me only after I added it.\"",
    "1735584": "The characteristics of this game are overfitting the training set. It is useful to add constraints and tricks at the feature level.\nUsing bn layer and feature loss is equivalent to assuming that the feature space, increases the features to distinguish, also eased the overfitting.",
    "1735647": "i have a feeling that normalisation layer is smiliar is the PCA trick before deep learning metric learning is \"invented\". \n\nIn traditional old ML methods, it is necessary to perform PCA to ensure that the distance is spherical in the embedding space. cosine distance is only meaningful if  the embedding space is spherical\n\nit is also a known trick that batch size should be large enough for metric learning/retrival/face recognition problem/contrastive learning",
    "1736096": "Is this for training or inference or both?",
    "1738614": "Yep, in [original paper](https://arxiv.org/pdf/1801.07698.pdf) they say, that:\n> After the last convolutional layer, we explore the BN-Dropout-FC-BN structure to get the final 512-D embedding feature.",
    "1741804": "Batch norm parameters are being updated during training\n\nIn the inference phase, the batch norm parameters called gamma and beta (which are learned in the training phase) and their running mean and std are fixed and will be used for computing the output.",
    "1750734": "Thank you for sharing great information!\nCould you show me some example of  kernels better than others, or give me hint to find better notebook.\nI can't judge what notebook is better or not...",
    "1750945": "Most of the time there is no wrong or right in machine learning. Try all options and see what works best.",
    "1751128": "Indeed. There's still a lot to be studied, so it's hard to say if something is right or wrong."
  },
  "source": "meta"
}