{
  "id": 406630,
  "title": "28th place solution - Using Arcface is not trivial.",
  "url": "/competitions/asl-signs/writeups/tmax-ai-28th-place-solution-using-arcface-is-not-t",
  "author_name": "",
  "post_date": "2023-12-26T00:37:25.563Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to Kaggle and Google for hosting the competition :)<br>\nAll of the work is done with <a href=\"https://www.kaggle.com/ydjoo12\" target=\"_blank\">@ydjoo12</a> </p>\n<h2>Summary</h2>\n<p>We used 5-layers Transformer encoder with cross entropy and sub-class Arcface loss and ensemble 6-models trained with different combintation of cross entropy and sub-class Arcface loss.</p>\n<h2>Data Preprocessing</h2>\n<ul>\n<li><p>129 keypoints are used. (21 for each hand, 40 for lip, 11 for pose, 16 for each eye, 4 for nose)</p></li>\n<li><p>Only use x and y.</p></li>\n<li><p>Pipeline: Augmentations, frame sampling, normalization, feature engineering and 0 imputations for NaN values.</p></li>\n<li><p>Frame sampling: When a length exceeds the maximum length, sampling was done at regular intervals up to the maximum length. It is better than center-crop in local CV and it minimized performance degradation while reducing maximum length.</p>\n<pre><code>L = (xyz)\n L &gt; max_len:\n    step = (L - ) // (max_len - )\n    indices = [i * step  i  (max_len)]\n    xyz = xyz[indices]\nL = (xyz)\n</code></pre></li>\n<li><p>Normalization: Standardization x and y independently.</p></li>\n<li><p>Feature Engineering</p>\n<ul>\n<li>motion: current xy - future xy</li>\n<li>hand joint distance</li>\n<li>time reverse difference: xy - xy.flip(dims=[0]): +0.004 on CV</li></ul></li>\n</ul>\n<h2>Augmentations</h2>\n<ul>\n<li><p>Flip pose: </p>\n<ul>\n<li>x_new = f(-(x - 0.5) + 0.5), where f is index change function.</li>\n<li>+0.01 on Public LB</li>\n<li>xy values are in [0, 1]. Shift to origin by subtracting 0.5 before flip, Back to original coordinates by adding 0.5</li>\n<li>There was no performance improvement when not moving to the origin.</li></ul></li>\n<li><p>Rotate pose: </p>\n<pre><code> ():\n    radian = np.radians(theta)\n    mat = np.array(\n        [[np.cos(radian), -np.sin(radian)], [np.sin(radian), np.cos(radian)]]\n    )\n    xyz[:, :, :] = xyz[:, :, :] - \n    xyz_reshape = xyz.reshape(-, )\n    xyz_rotate = np.dot(xyz_reshape, mat).reshape(xyz.shape)\n     xyz_rotate[:, :, :] + \n</code></pre>\n<ul>\n<li>angle between -13 degree ~ 13 degree</li>\n<li><a href=\"https://openaccess.thecvf.com/content/WACV2022W/HADCV/papers/Bohacek_Sign_Pose-Based_Transformer_for_Word-Level_Sign_Language_Recognition_WACVW_2022_paper.pdf\" target=\"_blank\">Reference</a></li></ul></li>\n<li><p>Interpolation (Up and Down)</p>\n<pre><code>xyz = F.interpolate(\n        xyz, size=(V * C, resize), mode=, align_corners=\n).squeeze()\n</code></pre>\n<ul>\n<li>Up-sampling(Down-sampling) up to 125%(75%) of original length </li></ul></li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>5-layers Transformer encoder with weighted cross entropy<ul>\n<li>Class weight is based on performance of the class.</li>\n<li>Accuracy on LB and CV consistently improved until stacking 4~5 layers.</li>\n<li>Public LBs =&gt; 1-layer(5 seeds): 0.714, 2-layers(5 seeds): 0.738, 3-layers(5 seeds): 0.748, 4-layers(5 seeds): 0.751</li>\n<li>We add dropout layers after self-attention and fc layer based on the PyTorch official code.</li></ul></li>\n<li>5-layers Transformer encoder with weighted cross entropy and sub-class Arcface loss.<ul>\n<li>loss = cross entropy + 0.2 * subclass(K=3) Arcface</li>\n<li>loss = 0.2 * cross entropy + subclass(K=3) Arcface</li></ul></li>\n<li>Arcface<ul>\n<li>+0.01 on Public LB</li>\n<li>Using arcface loss alone resulted in worse performance.</li>\n<li>Arcface armed with cross entropy converges much faster and better than cross entropy alone.</li>\n<li>Subclass K=3, margin=0.2, scale=32</li></ul></li>\n<li>Scheduled Dropout<ul>\n<li>+0.002 on CV</li>\n<li>Dropout rate on final [CLS] increased x2 after half of training epochs</li></ul></li>\n<li>Label Smoothing<ul>\n<li>parameter: 0.2</li>\n<li>+0.01 on CV</li></ul></li>\n<li>Hyper parameters<ul>\n<li>Epochs: 140</li>\n<li>Max lenght: 64</li>\n<li>batch size: 64</li>\n<li>embed dim: 256</li>\n<li>num head: 4</li>\n<li>num layers: 5</li>\n<li>CosineAnnealingWarmRestarts w/ lr 1e-3 and AdamW</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>2 different seeds of Transformer with weighted cross entropy<ul>\n<li>Single model LB: 0.75</li></ul></li>\n<li>2 different seeds of weighted cross entropy + 0.2 * subclass(K=3) Arcface<ul>\n<li>Inference: Weighted ensemble of cross entropy and Arcface head.</li>\n<li>Single model LB: 0.76</li></ul></li>\n<li>2 different seeds of 0.2 * cross entropy + subclass(K=3) Arcface<ul>\n<li>Inference: Weighted ensemble of cross entropy and Arcface head.</li>\n<li>Single model LB: 0.75</li></ul></li>\n<li>6 models ensemble Public LB: 0.779</li>\n<li>All models are fp16</li>\n<li>Total Size: 20Mb</li>\n<li>Latency: 60ms/sample</li>\n</ul>\n<h2>Working on CV but not included in final submission</h2>\n<ul>\n<li>TTA<ul>\n<li>+0.000x on CV</li>\n<li>Sumbission Scoring Error. It might be a memory issue. </li></ul></li>\n<li>Angle between bones of each arm<ul>\n<li>0.000x on 1 fold. We couldn't fully validate it due to time limits.</li></ul></li>\n</ul>\n<h2>Not working</h2>\n<ul>\n<li>GCN embedding layer instead of Linear</li>\n<li>Stacking Spatial Attention &amp; Temporal Conv. blocks</li>\n<li>Distance between pose keypoints</li>\n<li>Removing outlier and Retraining<ul>\n<li>We used anlge between learned Arcface subclass vector</li>\n<li>About 5% (~4000 samples) are removed</li></ul></li>\n<li>Knowledge distillation with bigger Transformer</li>\n<li>Stochastic Weight Average</li>\n<li>Using all [CLS] in every layer</li>\n<li>Average all tokens instead of [CLS] token</li>\n<li>Stacking with MLP as meta-learner</li>\n</ul>",
  "messages": [
    {
      "id": "2243683",
      "postDate": "05/03/2023 04:56:18",
      "content": "<p>Thanks to Kaggle and Google for hosting the competition :)<br>\nAll of the work is done with <a href=\"https://www.kaggle.com/ydjoo12\" target=\"_blank\">@ydjoo12</a> </p>\n<h2>Summary</h2>\n<p>We used 5-layers Transformer encoder with cross entropy and sub-class Arcface loss and ensemble 6-models trained with different combintation of cross entropy and sub-class Arcface loss.</p>\n<h2>Data Preprocessing</h2>\n<ul>\n<li><p>129 keypoints are used. (21 for each hand, 40 for lip, 11 for pose, 16 for each eye, 4 for nose)</p></li>\n<li><p>Only use x and y.</p></li>\n<li><p>Pipeline: Augmentations, frame sampling, normalization, feature engineering and 0 imputations for NaN values.</p></li>\n<li><p>Frame sampling: When a length exceeds the maximum length, sampling was done at regular intervals up to the maximum length. It is better than center-crop in local CV and it minimized performance degradation while reducing maximum length.</p>\n<pre><code>L = (xyz)\n L &gt; max_len:\n    step = (L - ) // (max_len - )\n    indices = [i * step  i  (max_len)]\n    xyz = xyz[indices]\nL = (xyz)\n</code></pre></li>\n<li><p>Normalization: Standardization x and y independently.</p></li>\n<li><p>Feature Engineering</p>\n<ul>\n<li>motion: current xy - future xy</li>\n<li>hand joint distance</li>\n<li>time reverse difference: xy - xy.flip(dims=[0]): +0.004 on CV</li></ul></li>\n</ul>\n<h2>Augmentations</h2>\n<ul>\n<li><p>Flip pose: </p>\n<ul>\n<li>x_new = f(-(x - 0.5) + 0.5), where f is index change function.</li>\n<li>+0.01 on Public LB</li>\n<li>xy values are in [0, 1]. Shift to origin by subtracting 0.5 before flip, Back to original coordinates by adding 0.5</li>\n<li>There was no performance improvement when not moving to the origin.</li></ul></li>\n<li><p>Rotate pose: </p>\n<pre><code> ():\n    radian = np.radians(theta)\n    mat = np.array(\n        [[np.cos(radian), -np.sin(radian)], [np.sin(radian), np.cos(radian)]]\n    )\n    xyz[:, :, :] = xyz[:, :, :] - \n    xyz_reshape = xyz.reshape(-, )\n    xyz_rotate = np.dot(xyz_reshape, mat).reshape(xyz.shape)\n     xyz_rotate[:, :, :] + \n</code></pre>\n<ul>\n<li>angle between -13 degree ~ 13 degree</li>\n<li><a href=\"https://openaccess.thecvf.com/content/WACV2022W/HADCV/papers/Bohacek_Sign_Pose-Based_Transformer_for_Word-Level_Sign_Language_Recognition_WACVW_2022_paper.pdf\" target=\"_blank\">Reference</a></li></ul></li>\n<li><p>Interpolation (Up and Down)</p>\n<pre><code>xyz = F.interpolate(\n        xyz, size=(V * C, resize), mode=, align_corners=\n).squeeze()\n</code></pre>\n<ul>\n<li>Up-sampling(Down-sampling) up to 125%(75%) of original length </li></ul></li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>5-layers Transformer encoder with weighted cross entropy<ul>\n<li>Class weight is based on performance of the class.</li>\n<li>Accuracy on LB and CV consistently improved until stacking 4~5 layers.</li>\n<li>Public LBs =&gt; 1-layer(5 seeds): 0.714, 2-layers(5 seeds): 0.738, 3-layers(5 seeds): 0.748, 4-layers(5 seeds): 0.751</li>\n<li>We add dropout layers after self-attention and fc layer based on the PyTorch official code.</li></ul></li>\n<li>5-layers Transformer encoder with weighted cross entropy and sub-class Arcface loss.<ul>\n<li>loss = cross entropy + 0.2 * subclass(K=3) Arcface</li>\n<li>loss = 0.2 * cross entropy + subclass(K=3) Arcface</li></ul></li>\n<li>Arcface<ul>\n<li>+0.01 on Public LB</li>\n<li>Using arcface loss alone resulted in worse performance.</li>\n<li>Arcface armed with cross entropy converges much faster and better than cross entropy alone.</li>\n<li>Subclass K=3, margin=0.2, scale=32</li></ul></li>\n<li>Scheduled Dropout<ul>\n<li>+0.002 on CV</li>\n<li>Dropout rate on final [CLS] increased x2 after half of training epochs</li></ul></li>\n<li>Label Smoothing<ul>\n<li>parameter: 0.2</li>\n<li>+0.01 on CV</li></ul></li>\n<li>Hyper parameters<ul>\n<li>Epochs: 140</li>\n<li>Max lenght: 64</li>\n<li>batch size: 64</li>\n<li>embed dim: 256</li>\n<li>num head: 4</li>\n<li>num layers: 5</li>\n<li>CosineAnnealingWarmRestarts w/ lr 1e-3 and AdamW</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>2 different seeds of Transformer with weighted cross entropy<ul>\n<li>Single model LB: 0.75</li></ul></li>\n<li>2 different seeds of weighted cross entropy + 0.2 * subclass(K=3) Arcface<ul>\n<li>Inference: Weighted ensemble of cross entropy and Arcface head.</li>\n<li>Single model LB: 0.76</li></ul></li>\n<li>2 different seeds of 0.2 * cross entropy + subclass(K=3) Arcface<ul>\n<li>Inference: Weighted ensemble of cross entropy and Arcface head.</li>\n<li>Single model LB: 0.75</li></ul></li>\n<li>6 models ensemble Public LB: 0.779</li>\n<li>All models are fp16</li>\n<li>Total Size: 20Mb</li>\n<li>Latency: 60ms/sample</li>\n</ul>\n<h2>Working on CV but not included in final submission</h2>\n<ul>\n<li>TTA<ul>\n<li>+0.000x on CV</li>\n<li>Sumbission Scoring Error. It might be a memory issue. </li></ul></li>\n<li>Angle between bones of each arm<ul>\n<li>0.000x on 1 fold. We couldn't fully validate it due to time limits.</li></ul></li>\n</ul>\n<h2>Not working</h2>\n<ul>\n<li>GCN embedding layer instead of Linear</li>\n<li>Stacking Spatial Attention &amp; Temporal Conv. blocks</li>\n<li>Distance between pose keypoints</li>\n<li>Removing outlier and Retraining<ul>\n<li>We used anlge between learned Arcface subclass vector</li>\n<li>About 5% (~4000 samples) are removed</li></ul></li>\n<li>Knowledge distillation with bigger Transformer</li>\n<li>Stochastic Weight Average</li>\n<li>Using all [CLS] in every layer</li>\n<li>Average all tokens instead of [CLS] token</li>\n<li>Stacking with MLP as meta-learner</li>\n</ul>",
      "rawMarkdown": "Thanks to Kaggle and Google for hosting the competition :)\nAll of the work is done with @ydjoo12 \n\n## Summary\nWe used 5-layers Transformer encoder with cross entropy and sub-class Arcface loss and ensemble 6-models trained with different combintation of cross entropy and sub-class Arcface loss.\n\n## Data Preprocessing\n- 129 keypoints are used. (21 for each hand, 40 for lip, 11 for pose, 16 for each eye, 4 for nose)\n- Only use x and y.\n- Pipeline: Augmentations, frame sampling, normalization, feature engineering and 0 imputations for NaN values.\n- Frame sampling: When a length exceeds the maximum length, sampling was done at regular intervals up to the maximum length. It is better than center-crop in local CV and it minimized performance degradation while reducing maximum length.\n    ```python\n    L = len(xyz)\n    if L > max_len:\n        step = (L - 1) // (max_len - 1)\n        indices = [i * step for i in range(max_len)]\n        xyz = xyz[indices]\n    L = len(xyz)\n    ```\n  \n- Normalization: Standardization x and y independently.\n- Feature Engineering\n  - motion: current xy - future xy\n  - hand joint distance\n  - time reverse difference: xy - xy.flip(dims=[0]): +0.004 on CV\n\n\n## Augmentations\n- Flip pose: \n  - x_new = f(-(x - 0.5) + 0.5), where f is index change function.\n  - +0.01 on Public LB\n  - xy values are in [0, 1]. Shift to origin by subtracting 0.5 before flip, Back to original coordinates by adding 0.5\n  - There was no performance improvement when not moving to the origin.\n- Rotate pose: \n\n    ```python\n    def rotate(xyz, theta):\n        radian = np.radians(theta)\n        mat = np.array(\n            [[np.cos(radian), -np.sin(radian)], [np.sin(radian), np.cos(radian)]]\n        )\n        xyz[:, :, :2] = xyz[:, :, :2] - 0.5\n        xyz_reshape = xyz.reshape(-1, 2)\n        xyz_rotate = np.dot(xyz_reshape, mat).reshape(xyz.shape)\n        return xyz_rotate[:, :, :2] + 0.5\n    ```\n  - angle between -13 degree ~ 13 degree\n  - [Reference](https://openaccess.thecvf.com/content/WACV2022W/HADCV/papers/Bohacek_Sign_Pose-Based_Transformer_for_Word-Level_Sign_Language_Recognition_WACVW_2022_paper.pdf)\n- Interpolation (Up and Down)\n\n    ```python\n    xyz = F.interpolate(\n            xyz, size=(V * C, resize), mode=\"bilinear\", align_corners=False\n    ).squeeze()\n    ```\n  - Up-sampling(Down-sampling) up to 125%(75%) of original length \n\n\n## Training\n- 5-layers Transformer encoder with weighted cross entropy\n  - Class weight is based on performance of the class.\n  - Accuracy on LB and CV consistently improved until stacking 4~5 layers.\n    - Public LBs => 1-layer(5 seeds): 0.714, 2-layers(5 seeds): 0.738, 3-layers(5 seeds): 0.748, 4-layers(5 seeds): 0.751\n    - We add dropout layers after self-attention and fc layer based on the PyTorch official code.\n- 5-layers Transformer encoder with weighted cross entropy and sub-class Arcface loss.\n  - loss = cross entropy + 0.2 * subclass(K=3) Arcface\n  - loss = 0.2 * cross entropy + subclass(K=3) Arcface\n- Arcface\n  - +0.01 on Public LB\n  - Using arcface loss alone resulted in worse performance.\n  - Arcface armed with cross entropy converges much faster and better than cross entropy alone.\n  - Subclass K=3, margin=0.2, scale=32\n- Scheduled Dropout\n  - +0.002 on CV\n  - Dropout rate on final [CLS] increased x2 after half of training epochs\n- Label Smoothing\n  - parameter: 0.2\n  - +0.01 on CV\n- Hyper parameters\n  - Epochs: 140\n  - Max lenght: 64\n  - batch size: 64\n  - embed dim: 256\n  - num head: 4\n  - num layers: 5\n  - CosineAnnealingWarmRestarts w/ lr 1e-3 and AdamW\n\n\n## Ensemble\n- 2 different seeds of Transformer with weighted cross entropy\n  - Single model LB: 0.75\n- 2 different seeds of weighted cross entropy + 0.2 * subclass(K=3) Arcface\n  - Inference: Weighted ensemble of cross entropy and Arcface head.\n  - Single model LB: 0.76\n- 2 different seeds of 0.2 * cross entropy + subclass(K=3) Arcface\n  - Inference: Weighted ensemble of cross entropy and Arcface head.\n  - Single model LB: 0.75\n- 6 models ensemble Public LB: 0.779\n- All models are fp16\n- Total Size: 20Mb\n- Latency: 60ms/sample\n\n\n## Working on CV but not included in final submission\n- TTA\n  - +0.000x on CV\n  - Sumbission Scoring Error. It might be a memory issue. \n- Angle between bones of each arm\n  - 0.000x on 1 fold. We couldn't fully validate it due to time limits.\n\n## Not working\n- GCN embedding layer instead of Linear\n- Stacking Spatial Attention & Temporal Conv. blocks\n- Distance between pose keypoints\n- Removing outlier and Retraining\n  - We used anlge between learned Arcface subclass vector\n  - About 5% (~4000 samples) are removed\n- Knowledge distillation with bigger Transformer\n- Stochastic Weight Average\n- Using all [CLS] in every layer\n- Average all tokens instead of [CLS] token\n- Stacking with MLP as meta-learner",
      "votes": null
    },
    {
      "id": "2243738",
      "postDate": "05/03/2023 06:11:31",
      "content": "<p>I have noticed that Arcface is frequently used in top solutions of various competitions. <br>\nHowever, when I tried it myself, the convergence was very slow and the performance was not good at result. <br>\nIn my opinion, the randomly initialized class vectors may not be trained well. <br>\nDo you have any tips for using Arcface that you could share with me? <br>\n(e.g. Initialized class vector, Choosing margin and scale, normalization … )<br>\nIn my case, using Arcface in combination with cross entropy was effective.</p>",
      "rawMarkdown": "I have noticed that Arcface is frequently used in top solutions of various competitions. \nHowever, when I tried it myself, the convergence was very slow and the performance was not good at result. \nIn my opinion, the randomly initialized class vectors may not be trained well. \nDo you have any tips for using Arcface that you could share with me? \n(e.g. Initialized class vector, Choosing margin and scale, normalization ... )\nIn my case, using Arcface in combination with cross entropy was effective.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2243738,
      "author_name": "wonjunjg",
      "author_url": "",
      "post_date": "05/03/2023 06:11:31",
      "content": "<p>I have noticed that Arcface is frequently used in top solutions of various competitions. <br>\nHowever, when I tried it myself, the convergence was very slow and the performance was not good at result. <br>\nIn my opinion, the randomly initialized class vectors may not be trained well. <br>\nDo you have any tips for using Arcface that you could share with me? <br>\n(e.g. Initialized class vector, Choosing margin and scale, normalization … )<br>\nIn my case, using Arcface in combination with cross entropy was effective.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2243683": "Thanks to Kaggle and Google for hosting the competition :)\nAll of the work is done with @ydjoo12 \n\n## Summary\nWe used 5-layers Transformer encoder with cross entropy and sub-class Arcface loss and ensemble 6-models trained with different combintation of cross entropy and sub-class Arcface loss.\n\n## Data Preprocessing\n- 129 keypoints are used. (21 for each hand, 40 for lip, 11 for pose, 16 for each eye, 4 for nose)\n- Only use x and y.\n- Pipeline: Augmentations, frame sampling, normalization, feature engineering and 0 imputations for NaN values.\n- Frame sampling: When a length exceeds the maximum length, sampling was done at regular intervals up to the maximum length. It is better than center-crop in local CV and it minimized performance degradation while reducing maximum length.\n    ```python\n    L = len(xyz)\n    if L > max_len:\n        step = (L - 1) // (max_len - 1)\n        indices = [i * step for i in range(max_len)]\n        xyz = xyz[indices]\n    L = len(xyz)\n    ```\n  \n- Normalization: Standardization x and y independently.\n- Feature Engineering\n  - motion: current xy - future xy\n  - hand joint distance\n  - time reverse difference: xy - xy.flip(dims=[0]): +0.004 on CV\n\n\n## Augmentations\n- Flip pose: \n  - x_new = f(-(x - 0.5) + 0.5), where f is index change function.\n  - +0.01 on Public LB\n  - xy values are in [0, 1]. Shift to origin by subtracting 0.5 before flip, Back to original coordinates by adding 0.5\n  - There was no performance improvement when not moving to the origin.\n- Rotate pose: \n\n    ```python\n    def rotate(xyz, theta):\n        radian = np.radians(theta)\n        mat = np.array(\n            [[np.cos(radian), -np.sin(radian)], [np.sin(radian), np.cos(radian)]]\n        )\n        xyz[:, :, :2] = xyz[:, :, :2] - 0.5\n        xyz_reshape = xyz.reshape(-1, 2)\n        xyz_rotate = np.dot(xyz_reshape, mat).reshape(xyz.shape)\n        return xyz_rotate[:, :, :2] + 0.5\n    ```\n  - angle between -13 degree ~ 13 degree\n  - [Reference](https://openaccess.thecvf.com/content/WACV2022W/HADCV/papers/Bohacek_Sign_Pose-Based_Transformer_for_Word-Level_Sign_Language_Recognition_WACVW_2022_paper.pdf)\n- Interpolation (Up and Down)\n\n    ```python\n    xyz = F.interpolate(\n            xyz, size=(V * C, resize), mode=\"bilinear\", align_corners=False\n    ).squeeze()\n    ```\n  - Up-sampling(Down-sampling) up to 125%(75%) of original length \n\n\n## Training\n- 5-layers Transformer encoder with weighted cross entropy\n  - Class weight is based on performance of the class.\n  - Accuracy on LB and CV consistently improved until stacking 4~5 layers.\n    - Public LBs => 1-layer(5 seeds): 0.714, 2-layers(5 seeds): 0.738, 3-layers(5 seeds): 0.748, 4-layers(5 seeds): 0.751\n    - We add dropout layers after self-attention and fc layer based on the PyTorch official code.\n- 5-layers Transformer encoder with weighted cross entropy and sub-class Arcface loss.\n  - loss = cross entropy + 0.2 * subclass(K=3) Arcface\n  - loss = 0.2 * cross entropy + subclass(K=3) Arcface\n- Arcface\n  - +0.01 on Public LB\n  - Using arcface loss alone resulted in worse performance.\n  - Arcface armed with cross entropy converges much faster and better than cross entropy alone.\n  - Subclass K=3, margin=0.2, scale=32\n- Scheduled Dropout\n  - +0.002 on CV\n  - Dropout rate on final [CLS] increased x2 after half of training epochs\n- Label Smoothing\n  - parameter: 0.2\n  - +0.01 on CV\n- Hyper parameters\n  - Epochs: 140\n  - Max lenght: 64\n  - batch size: 64\n  - embed dim: 256\n  - num head: 4\n  - num layers: 5\n  - CosineAnnealingWarmRestarts w/ lr 1e-3 and AdamW\n\n\n## Ensemble\n- 2 different seeds of Transformer with weighted cross entropy\n  - Single model LB: 0.75\n- 2 different seeds of weighted cross entropy + 0.2 * subclass(K=3) Arcface\n  - Inference: Weighted ensemble of cross entropy and Arcface head.\n  - Single model LB: 0.76\n- 2 different seeds of 0.2 * cross entropy + subclass(K=3) Arcface\n  - Inference: Weighted ensemble of cross entropy and Arcface head.\n  - Single model LB: 0.75\n- 6 models ensemble Public LB: 0.779\n- All models are fp16\n- Total Size: 20Mb\n- Latency: 60ms/sample\n\n\n## Working on CV but not included in final submission\n- TTA\n  - +0.000x on CV\n  - Sumbission Scoring Error. It might be a memory issue. \n- Angle between bones of each arm\n  - 0.000x on 1 fold. We couldn't fully validate it due to time limits.\n\n## Not working\n- GCN embedding layer instead of Linear\n- Stacking Spatial Attention & Temporal Conv. blocks\n- Distance between pose keypoints\n- Removing outlier and Retraining\n  - We used anlge between learned Arcface subclass vector\n  - About 5% (~4000 samples) are removed\n- Knowledge distillation with bigger Transformer\n- Stochastic Weight Average\n- Using all [CLS] in every layer\n- Average all tokens instead of [CLS] token\n- Stacking with MLP as meta-learner",
    "2243738": "I have noticed that Arcface is frequently used in top solutions of various competitions. \nHowever, when I tried it myself, the convergence was very slow and the performance was not good at result. \nIn my opinion, the randomly initialized class vectors may not be trained well. \nDo you have any tips for using Arcface that you could share with me? \n(e.g. Initialized class vector, Choosing margin and scale, normalization ... )\nIn my case, using Arcface in combination with cross entropy was effective."
  },
  "source": "meta"
}