{
  "id": 243951,
  "title": "15th adversarial validation etc",
  "url": "/competitions/bms-molecular-translation/writeups/akirasosa-15th-adversarial-validation-etc",
  "author_name": "",
  "post_date": "2021-06-04T16:21:14.943Z",
  "votes": 22,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, I would like to say thanks to Kaggle and organizers hosts this exciting competition. Congrats winners who won this tough competition! </p>\n<h2>Findings</h2>\n<p>I have tried adversarial validation in early stage. Here is the result.</p>\n<p><img src=\"https://ucaa060899a3c8e0931bf9d91ade.previews.dropboxusercontent.com/p/thumb/ABLhfY6YZmPuwcl_I1J-i3lYgYixcNXnWlOyDMBIigA4YS88S8InmWe-QhG8Y2MoClDpnXxH96ia6Cy38-rarws-WL8pB59Q_DYay85J19n8R-jp-ucGx5cnkI4rMpbyGcmSO3aQGaPHoxMrdtEQuDchv_nJO8MoZRDJYXADTLCbDcptwZ806rIePExIe6-klx1th-v3kSfz267TdsYhG-YeJziUYy8FzTXr9xVn7gvaEDu40-REAbdxA_CqnyulEUq-8OS8kzBXv3rUs3kFqP9exr3AZ_0LRmFrJsN0Adbj4DJnBm_WUuZPWoXITU2OoG4E7sh1fekri7_ei0DSctsjzIuC_pJkOL7XouqLmmUd2NtWBflkqG5B6JLwK12yGzthW35ZOKL8R2hGZMJwezHG/p.png?fv_content=true&amp;size_mode=5\" alt=\"adv val\"></p>\n<p>This figure by itself tells how this test set was made. Test set can be split into two subsets. I call them as …</p>\n<ul>\n<li>Test subset A (left samples on figure which is not similar with train set)</li>\n<li>Test subset B (right samples on figure which is similar with train set)</li>\n</ul>\n<p>Then, I was pretty sure that pseudo labeling will work well for test set. I also thought that it's necessary to create models which fits both subset well. It's possible to estimate how much current model fits for each samples, if erase all half predictions.</p>\n<h2>Training and modeling</h2>\n<p>I used 3 models, Deit tiny, small and base.</p>\n<p>My model mainly predicts smiles instead of inchi. smiles is more compact than inchi. It helps reducing training time. smiles is converted inchi before submission. I also created a model to predict inchi, but it's for fallback when smiles is invalid. As a summary, I predicted inchi from 3 ways. The priority order is followings.</p>\n<ul>\n<li>norm inchi from predicted smiles</li>\n<li>norm inchi from predicted inchi</li>\n<li>raw predicted inchi</li>\n</ul>\n<p>Training steps consist from 3 steps.</p>\n<ol>\n<li>Pre-train only decoder by using randomized smiles canonicalization.</li>\n<li>Combine Deit with the decoder trained on step 1 and fine tune by extra synthesized images (size 224).</li>\n<li>Fine tune with train set (only hard samples) and pseudo labels with image size=384.</li>\n</ol>\n<p>At step1, I train transformer to predict canonicalized inchi and smiles from randomized smiles. Next, self attention and heads on decoder were transferred and combined with Deit. So, only cross attention on decoder is trained from scratch. As I don't have so many GPUs, 224 images are used at this step.</p>\n<p>Finally, I have trained 6 models at step 3.</p>\n<ul>\n<li>a1) Deit tiny trained by pseudo labels from Test subset A + train set (only hard samples).</li>\n<li>a2) Deit small trained by pseudo labels from Test subset A + train set (only hard samples).</li>\n<li>a3) Deit base trained by pseudo labels from Test subset A + train set (only hard samples).</li>\n<li>b1) Deit tiny trained by pseudo labels from Test subset B + train set (only hard samples).</li>\n<li>b2) Deit small trained by pseudo labels from Test subset B + train set (only hard samples).</li>\n<li>b3) Deit base trained by pseudo labels from Test subset B + train set (only hard samples).</li>\n</ul>\n<p>Position embedding was expanded for image size 384 by interpolation at step 3.</p>\n<h2>Inference</h2>\n<p>I used TTA (rot90) and logits averaging per token.</p>\n<p>For Test subset A, I used averaging weight 9:9:9:1:1:1 for models a1, a2, a3, b1, b2 and b3. On the other hand, for Test subset B, I used weight 1:1:1:9:9:9. It ends up with 0.64 on LB.</p>",
  "messages": [
    {
      "id": "1336017",
      "postDate": "06/04/2021 15:30:40",
      "content": "<p>First of all, I would like to say thanks to Kaggle and organizers hosts this exciting competition. Congrats winners who won this tough competition! </p>\n<h2>Findings</h2>\n<p>I have tried adversarial validation in early stage. Here is the result.</p>\n<p><img src=\"https://ucaa060899a3c8e0931bf9d91ade.previews.dropboxusercontent.com/p/thumb/ABLhfY6YZmPuwcl_I1J-i3lYgYixcNXnWlOyDMBIigA4YS88S8InmWe-QhG8Y2MoClDpnXxH96ia6Cy38-rarws-WL8pB59Q_DYay85J19n8R-jp-ucGx5cnkI4rMpbyGcmSO3aQGaPHoxMrdtEQuDchv_nJO8MoZRDJYXADTLCbDcptwZ806rIePExIe6-klx1th-v3kSfz267TdsYhG-YeJziUYy8FzTXr9xVn7gvaEDu40-REAbdxA_CqnyulEUq-8OS8kzBXv3rUs3kFqP9exr3AZ_0LRmFrJsN0Adbj4DJnBm_WUuZPWoXITU2OoG4E7sh1fekri7_ei0DSctsjzIuC_pJkOL7XouqLmmUd2NtWBflkqG5B6JLwK12yGzthW35ZOKL8R2hGZMJwezHG/p.png?fv_content=true&amp;size_mode=5\" alt=\"adv val\"></p>\n<p>This figure by itself tells how this test set was made. Test set can be split into two subsets. I call them as …</p>\n<ul>\n<li>Test subset A (left samples on figure which is not similar with train set)</li>\n<li>Test subset B (right samples on figure which is similar with train set)</li>\n</ul>\n<p>Then, I was pretty sure that pseudo labeling will work well for test set. I also thought that it's necessary to create models which fits both subset well. It's possible to estimate how much current model fits for each samples, if erase all half predictions.</p>\n<h2>Training and modeling</h2>\n<p>I used 3 models, Deit tiny, small and base.</p>\n<p>My model mainly predicts smiles instead of inchi. smiles is more compact than inchi. It helps reducing training time. smiles is converted inchi before submission. I also created a model to predict inchi, but it's for fallback when smiles is invalid. As a summary, I predicted inchi from 3 ways. The priority order is followings.</p>\n<ul>\n<li>norm inchi from predicted smiles</li>\n<li>norm inchi from predicted inchi</li>\n<li>raw predicted inchi</li>\n</ul>\n<p>Training steps consist from 3 steps.</p>\n<ol>\n<li>Pre-train only decoder by using randomized smiles canonicalization.</li>\n<li>Combine Deit with the decoder trained on step 1 and fine tune by extra synthesized images (size 224).</li>\n<li>Fine tune with train set (only hard samples) and pseudo labels with image size=384.</li>\n</ol>\n<p>At step1, I train transformer to predict canonicalized inchi and smiles from randomized smiles. Next, self attention and heads on decoder were transferred and combined with Deit. So, only cross attention on decoder is trained from scratch. As I don't have so many GPUs, 224 images are used at this step.</p>\n<p>Finally, I have trained 6 models at step 3.</p>\n<ul>\n<li>a1) Deit tiny trained by pseudo labels from Test subset A + train set (only hard samples).</li>\n<li>a2) Deit small trained by pseudo labels from Test subset A + train set (only hard samples).</li>\n<li>a3) Deit base trained by pseudo labels from Test subset A + train set (only hard samples).</li>\n<li>b1) Deit tiny trained by pseudo labels from Test subset B + train set (only hard samples).</li>\n<li>b2) Deit small trained by pseudo labels from Test subset B + train set (only hard samples).</li>\n<li>b3) Deit base trained by pseudo labels from Test subset B + train set (only hard samples).</li>\n</ul>\n<p>Position embedding was expanded for image size 384 by interpolation at step 3.</p>\n<h2>Inference</h2>\n<p>I used TTA (rot90) and logits averaging per token.</p>\n<p>For Test subset A, I used averaging weight 9:9:9:1:1:1 for models a1, a2, a3, b1, b2 and b3. On the other hand, for Test subset B, I used weight 1:1:1:9:9:9. It ends up with 0.64 on LB.</p>",
      "rawMarkdown": "First of all, I would like to say thanks to Kaggle and organizers hosts this exciting competition. Congrats winners who won this tough competition! \n\n## Findings\n\nI have tried adversarial validation in early stage. Here is the result.\n\n![adv val](https://ucaa060899a3c8e0931bf9d91ade.previews.dropboxusercontent.com/p/thumb/ABLhfY6YZmPuwcl_I1J-i3lYgYixcNXnWlOyDMBIigA4YS88S8InmWe-QhG8Y2MoClDpnXxH96ia6Cy38-rarws-WL8pB59Q_DYay85J19n8R-jp-ucGx5cnkI4rMpbyGcmSO3aQGaPHoxMrdtEQuDchv_nJO8MoZRDJYXADTLCbDcptwZ806rIePExIe6-klx1th-v3kSfz267TdsYhG-YeJziUYy8FzTXr9xVn7gvaEDu40-REAbdxA_CqnyulEUq-8OS8kzBXv3rUs3kFqP9exr3AZ_0LRmFrJsN0Adbj4DJnBm_WUuZPWoXITU2OoG4E7sh1fekri7_ei0DSctsjzIuC_pJkOL7XouqLmmUd2NtWBflkqG5B6JLwK12yGzthW35ZOKL8R2hGZMJwezHG/p.png?fv_content=true&size_mode=5)\n\nThis figure by itself tells how this test set was made. Test set can be split into two subsets. I call them as ...\n\n* Test subset A (left samples on figure which is not similar with train set)\n* Test subset B (right samples on figure which is similar with train set)\n\nThen, I was pretty sure that pseudo labeling will work well for test set. I also thought that it's necessary to create models which fits both subset well. It's possible to estimate how much current model fits for each samples, if erase all half predictions.\n\n\n## Training and modeling\n\nI used 3 models, Deit tiny, small and base.\n\nMy model mainly predicts smiles instead of inchi. smiles is more compact than inchi. It helps reducing training time. smiles is converted inchi before submission. I also created a model to predict inchi, but it's for fallback when smiles is invalid. As a summary, I predicted inchi from 3 ways. The priority order is followings.\n\n* norm inchi from predicted smiles\n* norm inchi from predicted inchi\n* raw predicted inchi\n\nTraining steps consist from 3 steps.\n\n1. Pre-train only decoder by using randomized smiles canonicalization.\n2. Combine Deit with the decoder trained on step 1 and fine tune by extra synthesized images (size 224).\n3. Fine tune with train set (only hard samples) and pseudo labels with image size=384.\n\nAt step1, I train transformer to predict canonicalized inchi and smiles from randomized smiles. Next, self attention and heads on decoder were transferred and combined with Deit. So, only cross attention on decoder is trained from scratch. As I don't have so many GPUs, 224 images are used at this step.\n\nFinally, I have trained 6 models at step 3.\n\n* a1) Deit tiny trained by pseudo labels from Test subset A + train set (only hard samples).\n* a2) Deit small trained by pseudo labels from Test subset A + train set (only hard samples).\n* a3) Deit base trained by pseudo labels from Test subset A + train set (only hard samples).\n* b1) Deit tiny trained by pseudo labels from Test subset B + train set (only hard samples).\n* b2) Deit small trained by pseudo labels from Test subset B + train set (only hard samples).\n* b3) Deit base trained by pseudo labels from Test subset B + train set (only hard samples).\n\nPosition embedding was expanded for image size 384 by interpolation at step 3.\n\n## Inference\n\nI used TTA (rot90) and logits averaging per token.\n\nFor Test subset A, I used averaging weight 9:9:9:1:1:1 for models a1, a2, a3, b1, b2 and b3. On the other hand, for Test subset B, I used weight 1:1:1:9:9:9. It ends up with 0.64 on LB.",
      "votes": null
    },
    {
      "id": "1336051",
      "postDate": "06/04/2021 15:51:35",
      "content": "<p>Great work! LB 0.65 as a solo competitor is amazing.<br>\nLet's fight together next time :)</p>",
      "rawMarkdown": "Great work! LB 0.65 as a solo competitor is amazing.\nLet's fight together next time :)",
      "votes": null
    },
    {
      "id": "1336497",
      "postDate": "06/05/2021 01:43:07",
      "content": "<p>Deit.   learned new staff.</p>",
      "rawMarkdown": "Deit.   learned new staff.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1336051,
      "author_name": "kfujikawa",
      "author_url": "",
      "post_date": "06/04/2021 15:51:35",
      "content": "<p>Great work! LB 0.65 as a solo competitor is amazing.<br>\nLet's fight together next time :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336497,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "06/05/2021 01:43:07",
      "content": "<p>Deit.   learned new staff.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1336017": "First of all, I would like to say thanks to Kaggle and organizers hosts this exciting competition. Congrats winners who won this tough competition! \n\n## Findings\n\nI have tried adversarial validation in early stage. Here is the result.\n\n![adv val](https://ucaa060899a3c8e0931bf9d91ade.previews.dropboxusercontent.com/p/thumb/ABLhfY6YZmPuwcl_I1J-i3lYgYixcNXnWlOyDMBIigA4YS88S8InmWe-QhG8Y2MoClDpnXxH96ia6Cy38-rarws-WL8pB59Q_DYay85J19n8R-jp-ucGx5cnkI4rMpbyGcmSO3aQGaPHoxMrdtEQuDchv_nJO8MoZRDJYXADTLCbDcptwZ806rIePExIe6-klx1th-v3kSfz267TdsYhG-YeJziUYy8FzTXr9xVn7gvaEDu40-REAbdxA_CqnyulEUq-8OS8kzBXv3rUs3kFqP9exr3AZ_0LRmFrJsN0Adbj4DJnBm_WUuZPWoXITU2OoG4E7sh1fekri7_ei0DSctsjzIuC_pJkOL7XouqLmmUd2NtWBflkqG5B6JLwK12yGzthW35ZOKL8R2hGZMJwezHG/p.png?fv_content=true&size_mode=5)\n\nThis figure by itself tells how this test set was made. Test set can be split into two subsets. I call them as ...\n\n* Test subset A (left samples on figure which is not similar with train set)\n* Test subset B (right samples on figure which is similar with train set)\n\nThen, I was pretty sure that pseudo labeling will work well for test set. I also thought that it's necessary to create models which fits both subset well. It's possible to estimate how much current model fits for each samples, if erase all half predictions.\n\n\n## Training and modeling\n\nI used 3 models, Deit tiny, small and base.\n\nMy model mainly predicts smiles instead of inchi. smiles is more compact than inchi. It helps reducing training time. smiles is converted inchi before submission. I also created a model to predict inchi, but it's for fallback when smiles is invalid. As a summary, I predicted inchi from 3 ways. The priority order is followings.\n\n* norm inchi from predicted smiles\n* norm inchi from predicted inchi\n* raw predicted inchi\n\nTraining steps consist from 3 steps.\n\n1. Pre-train only decoder by using randomized smiles canonicalization.\n2. Combine Deit with the decoder trained on step 1 and fine tune by extra synthesized images (size 224).\n3. Fine tune with train set (only hard samples) and pseudo labels with image size=384.\n\nAt step1, I train transformer to predict canonicalized inchi and smiles from randomized smiles. Next, self attention and heads on decoder were transferred and combined with Deit. So, only cross attention on decoder is trained from scratch. As I don't have so many GPUs, 224 images are used at this step.\n\nFinally, I have trained 6 models at step 3.\n\n* a1) Deit tiny trained by pseudo labels from Test subset A + train set (only hard samples).\n* a2) Deit small trained by pseudo labels from Test subset A + train set (only hard samples).\n* a3) Deit base trained by pseudo labels from Test subset A + train set (only hard samples).\n* b1) Deit tiny trained by pseudo labels from Test subset B + train set (only hard samples).\n* b2) Deit small trained by pseudo labels from Test subset B + train set (only hard samples).\n* b3) Deit base trained by pseudo labels from Test subset B + train set (only hard samples).\n\nPosition embedding was expanded for image size 384 by interpolation at step 3.\n\n## Inference\n\nI used TTA (rot90) and logits averaging per token.\n\nFor Test subset A, I used averaging weight 9:9:9:1:1:1 for models a1, a2, a3, b1, b2 and b3. On the other hand, for Test subset B, I used weight 1:1:1:9:9:9. It ends up with 0.64 on LB.",
    "1336051": "Great work! LB 0.65 as a solo competitor is amazing.\nLet's fight together next time :)",
    "1336497": "Deit.   learned new staff."
  },
  "source": "meta"
}