{
  "id": 80718,
  "title": "10th place solution - Meta embedding, EMA, Ensemble",
  "url": "/competitions/quora-insincere-questions-classification/writeups/tks-10th-place-solution-meta-embedding-ema-ensembl",
  "author_name": "",
  "post_date": "2019-02-16T14:33:01.537Z",
  "votes": 26,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Here are all my submissions. 1: keras, 2-4:pytorch</p>\n\n<ol>\n<li><p><a href=\"https://www.kaggle.com/tks0123456789/projection-meta-embedding-and-ema?scriptVersionId=7916644\">Projection meta embedding and EMA(version 1/4)</a>, 0.69480(Public)</p></li>\n<li><p><a href=\"https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10163202\">PME_EMA 6 x 8 pochs(version 2/14)</a>, 0.69568</p></li>\n<li><p><a href=\"https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10224276\">PME_EMA 6 x 8 pochs(version 10/14)</a>, 0.70551, 0.70964(Private)</p></li>\n<li><p><a href=\"https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10275816\">PME_EMA 6 x 8 pochs(version 14/14)</a>, 0.70061, 0.70921</p></li>\n</ol>\n\n<h2>Preprocessing</h2>\n\n<p>Separating punctuations only. Spell correction didn't work for me.</p>\n\n<h2>Model structure</h2>\n\n<p><strong>Average</strong> ensemble of 6 models of the same network.</p>\n\n<pre><code>Embedding(max_features, 600)\nLinear(in_features=600, out_features=128, bias=True)\nReLU()\nGRU(128, 128, batch_first=True, bidirectional=True)\nGlobalMaxPooling1D()\nLinear(in_features=256, out_features=256, bias=True)\nReLU()\nLinear(in_features=256, out_features=1, bias=True)\n</code></pre>\n\n<h2>Projection Meta Embedding(PME)</h2>\n\n<p>Meta embedding is a method for combining multiple pretrained embeddings and dicussed in <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71778\">3 Methods to combine embeddings</a>.\nPME is <strong>Unweighted DME ([4])+ ReLU</strong>. It concats several frozen pretrained embeddings and project to a lower dimensional space by linear layer with ReLU activation.</p>\n\n<h2>Exponential Moving Averaging of weights(EMA)</h2>\n\n<p>It caluculates exponential moving average of weights during training. It is usually done on a minibatch level. I chose 10 updates per epoch for speed.\n[1], [2], [3] use EMA.</p>\n\n<h2>Tuning for ensemble</h2>\n\n<h3>n_embed: #of pretrained embeddings in a single model.</h3>\n\n<p>I tried n_embed=1, 2, 3, 4, and 2 is best. Using 4 embeddings is better for a single model F1, but lacks model diversity, which cause worse ensemble performance.</p>\n\n<h3>epoch</h3>\n\n<p>The followings are mean F1 of 10-fold CV. The best epoch is 5 for a single model and 8 for ensemble.\n<img src=\"https://raw.githubusercontent.com/tks0123456789/kaggle-Quora/master/cv10.png\" alt=\"Image\"></p>\n\n<p>[1] <a href=\"https://arxiv.org/abs/1703.01780\">A. Tarvainen and H. Valpola(2017) Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results.</a></p>\n\n<p>[2] <a href=\"https://arxiv.org/abs/1804.09541\">Yu, A. W., Dohan, D., Luong, M.-T., Zhao, R., Chen, K., Norouzi, M., and Le, Q. V.(2018) QANet: Combining local convolution with global self-attention for reading comprehension</a></p>\n\n<p>[3] <a href=\"https://arxiv.org/abs/1611.01603\">Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi (2016) Bidirectional attention flow for machine comprehension.</a></p>\n\n<p>[4] <a href=\"https://arxiv.org/abs/1804.07983\">Douwe Kiela, Changhan Wang, Kyunghyun Cho (2018) Dynamic Meta-Embeddings for Improved Sentence Representations</a></p>",
  "messages": [
    {
      "id": "472364",
      "postDate": "02/15/2019 19:12:32",
      "content": "<p>Here are all my submissions. 1: keras, 2-4:pytorch</p>\n\n<ol>\n<li><p><a href=\"https://www.kaggle.com/tks0123456789/projection-meta-embedding-and-ema?scriptVersionId=7916644\">Projection meta embedding and EMA(version 1/4)</a>, 0.69480(Public)</p></li>\n<li><p><a href=\"https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10163202\">PME_EMA 6 x 8 pochs(version 2/14)</a>, 0.69568</p></li>\n<li><p><a href=\"https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10224276\">PME_EMA 6 x 8 pochs(version 10/14)</a>, 0.70551, 0.70964(Private)</p></li>\n<li><p><a href=\"https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10275816\">PME_EMA 6 x 8 pochs(version 14/14)</a>, 0.70061, 0.70921</p></li>\n</ol>\n\n<h2>Preprocessing</h2>\n\n<p>Separating punctuations only. Spell correction didn't work for me.</p>\n\n<h2>Model structure</h2>\n\n<p><strong>Average</strong> ensemble of 6 models of the same network.</p>\n\n<pre><code>Embedding(max_features, 600)\nLinear(in_features=600, out_features=128, bias=True)\nReLU()\nGRU(128, 128, batch_first=True, bidirectional=True)\nGlobalMaxPooling1D()\nLinear(in_features=256, out_features=256, bias=True)\nReLU()\nLinear(in_features=256, out_features=1, bias=True)\n</code></pre>\n\n<h2>Projection Meta Embedding(PME)</h2>\n\n<p>Meta embedding is a method for combining multiple pretrained embeddings and dicussed in <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71778\">3 Methods to combine embeddings</a>.\nPME is <strong>Unweighted DME ([4])+ ReLU</strong>. It concats several frozen pretrained embeddings and project to a lower dimensional space by linear layer with ReLU activation.</p>\n\n<h2>Exponential Moving Averaging of weights(EMA)</h2>\n\n<p>It caluculates exponential moving average of weights during training. It is usually done on a minibatch level. I chose 10 updates per epoch for speed.\n[1], [2], [3] use EMA.</p>\n\n<h2>Tuning for ensemble</h2>\n\n<h3>n_embed: #of pretrained embeddings in a single model.</h3>\n\n<p>I tried n_embed=1, 2, 3, 4, and 2 is best. Using 4 embeddings is better for a single model F1, but lacks model diversity, which cause worse ensemble performance.</p>\n\n<h3>epoch</h3>\n\n<p>The followings are mean F1 of 10-fold CV. The best epoch is 5 for a single model and 8 for ensemble.\n<img src=\"https://raw.githubusercontent.com/tks0123456789/kaggle-Quora/master/cv10.png\" alt=\"Image\"></p>\n\n<p>[1] <a href=\"https://arxiv.org/abs/1703.01780\">A. Tarvainen and H. Valpola(2017) Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results.</a></p>\n\n<p>[2] <a href=\"https://arxiv.org/abs/1804.09541\">Yu, A. W., Dohan, D., Luong, M.-T., Zhao, R., Chen, K., Norouzi, M., and Le, Q. V.(2018) QANet: Combining local convolution with global self-attention for reading comprehension</a></p>\n\n<p>[3] <a href=\"https://arxiv.org/abs/1611.01603\">Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi (2016) Bidirectional attention flow for machine comprehension.</a></p>\n\n<p>[4] <a href=\"https://arxiv.org/abs/1804.07983\">Douwe Kiela, Changhan Wang, Kyunghyun Cho (2018) Dynamic Meta-Embeddings for Improved Sentence Representations</a></p>",
      "rawMarkdown": "Here are all my submissions. 1: keras, 2-4:pytorch\n\n1. [Projection meta embedding and EMA(version 1/4)](https://www.kaggle.com/tks0123456789/projection-meta-embedding-and-ema?scriptVersionId=7916644), 0.69480(Public)\n\n2. [PME_EMA 6 x 8 pochs(version 2/14)](https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10163202), 0.69568\n3. [PME_EMA 6 x 8 pochs(version 10/14)](https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10224276), 0.70551, 0.70964(Private)\n\n4. [PME_EMA 6 x 8 pochs(version 14/14)](https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10275816), 0.70061, 0.70921\n\n## Preprocessing\nSeparating punctuations only. Spell correction didn't work for me.\n\n## Model structure\n**Average** ensemble of 6 models of the same network.\n\n    Embedding(max_features, 600)\n    Linear(in_features=600, out_features=128, bias=True)\n    ReLU()\n    GRU(128, 128, batch_first=True, bidirectional=True)\n    GlobalMaxPooling1D()\n    Linear(in_features=256, out_features=256, bias=True)\n    ReLU()\n    Linear(in_features=256, out_features=1, bias=True)\n\n## Projection Meta Embedding(PME) \n\nMeta embedding is a method for combining multiple pretrained embeddings and dicussed in [3 Methods to combine embeddings](https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71778).\nPME is **Unweighted DME ([4])+ ReLU**. It concats several frozen pretrained embeddings and project to a lower dimensional space by linear layer with ReLU activation.\n\n\n\n## Exponential Moving Averaging of weights(EMA)\n\nIt caluculates exponential moving average of weights during training. It is usually done on a minibatch level. I chose 10 updates per epoch for speed.\n[1], [2], [3] use EMA.\n\n\n## Tuning for ensemble\n\n### n_embed: #of pretrained embeddings in a single model.\nI tried n_embed=1, 2, 3, 4, and 2 is best. Using 4 embeddings is better for a single model F1, but lacks model diversity, which cause worse ensemble performance.\n\n### epoch\nThe followings are mean F1 of 10-fold CV. The best epoch is 5 for a single model and 8 for ensemble.\n![Image](https://raw.githubusercontent.com/tks0123456789/kaggle-Quora/master/cv10.png)\n\n\n[1] [A. Tarvainen and H. Valpola(2017) Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results.](https://arxiv.org/abs/1703.01780)\n\n[2] [Yu, A. W., Dohan, D., Luong, M.-T., Zhao, R., Chen, K., Norouzi, M., and Le, Q. V.(2018) QANet: Combining local convolution with global self-attention for reading comprehension](https://arxiv.org/abs/1804.09541)\n\n[3] [Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi (2016) Bidirectional attention flow for machine comprehension.](https://arxiv.org/abs/1611.01603)\n\n[4] [Douwe Kiela, Changhan Wang, Kyunghyun Cho (2018) Dynamic Meta-Embeddings for Improved Sentence Representations](https://arxiv.org/abs/1804.07983)",
      "votes": null
    },
    {
      "id": "472502",
      "postDate": "02/16/2019 03:49:31",
      "content": "<p>Great solution! I mean \"great\" refers to difference from common shared kernel, like EMA, embedding combination variety for model ensembles. But I have several questions. Do you have an idea about how many improvements brought by PME and EMA? I have looked at your codes. I do not find the 10-fold cv part. It seems you do not do cross validation but how do you evaluate your model performance and reach the good balance between overfitting and underfitting? Do you publish the kernel that generates the 10-fold cv model performance plot against epoch?</p>",
      "rawMarkdown": "Great solution! I mean \"great\" refers to difference from common shared kernel, like EMA, embedding combination variety for model ensembles. But I have several questions. Do you have an idea about how many improvements brought by PME and EMA? I have looked at your codes. I do not find the 10-fold cv part. It seems you do not do cross validation but how do you evaluate your model performance and reach the good balance between overfitting and underfitting? Do you publish the kernel that generates the 10-fold cv model performance plot against epoch?",
      "votes": null
    },
    {
      "id": "472765",
      "postDate": "02/16/2019 16:28:55",
      "content": "<p>Congratulations @tks </p>",
      "rawMarkdown": "Congratulations @tks",
      "votes": null
    },
    {
      "id": "473361",
      "postDate": "02/17/2019 21:08:40",
      "content": "<p>0.0015-0.002 for EMA. I don't know improvement by PME because I didn't try other embedding method. Code and result are <a href=\"https://github.com/tks0123456789/kaggle-Quora\">here</a>.</p>",
      "rawMarkdown": "0.0015-0.002 for EMA. I don't know improvement by PME because I didn't try other embedding method. Code and result are [here](https://github.com/tks0123456789/kaggle-Quora).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 472502,
      "author_name": "nzholmes",
      "author_url": "",
      "post_date": "02/16/2019 03:49:31",
      "content": "<p>Great solution! I mean \"great\" refers to difference from common shared kernel, like EMA, embedding combination variety for model ensembles. But I have several questions. Do you have an idea about how many improvements brought by PME and EMA? I have looked at your codes. I do not find the 10-fold cv part. It seems you do not do cross validation but how do you evaluate your model performance and reach the good balance between overfitting and underfitting? Do you publish the kernel that generates the 10-fold cv model performance plot against epoch?</p>",
      "votes": null,
      "replies": [
        {
          "id": 473361,
          "author_name": "tks0123456789",
          "author_url": "",
          "post_date": "02/17/2019 21:08:40",
          "content": "<p>0.0015-0.002 for EMA. I don't know improvement by PME because I didn't try other embedding method. Code and result are <a href=\"https://github.com/tks0123456789/kaggle-Quora\">here</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472765,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "02/16/2019 16:28:55",
      "content": "<p>Congratulations @tks </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "472364": "Here are all my submissions. 1: keras, 2-4:pytorch\n\n1. [Projection meta embedding and EMA(version 1/4)](https://www.kaggle.com/tks0123456789/projection-meta-embedding-and-ema?scriptVersionId=7916644), 0.69480(Public)\n\n2. [PME_EMA 6 x 8 pochs(version 2/14)](https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10163202), 0.69568\n3. [PME_EMA 6 x 8 pochs(version 10/14)](https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10224276), 0.70551, 0.70964(Private)\n\n4. [PME_EMA 6 x 8 pochs(version 14/14)](https://www.kaggle.com/tks0123456789/pme-ema-6-x-8-pochs?scriptVersionId=10275816), 0.70061, 0.70921\n\n## Preprocessing\nSeparating punctuations only. Spell correction didn't work for me.\n\n## Model structure\n**Average** ensemble of 6 models of the same network.\n\n    Embedding(max_features, 600)\n    Linear(in_features=600, out_features=128, bias=True)\n    ReLU()\n    GRU(128, 128, batch_first=True, bidirectional=True)\n    GlobalMaxPooling1D()\n    Linear(in_features=256, out_features=256, bias=True)\n    ReLU()\n    Linear(in_features=256, out_features=1, bias=True)\n\n## Projection Meta Embedding(PME) \n\nMeta embedding is a method for combining multiple pretrained embeddings and dicussed in [3 Methods to combine embeddings](https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71778).\nPME is **Unweighted DME ([4])+ ReLU**. It concats several frozen pretrained embeddings and project to a lower dimensional space by linear layer with ReLU activation.\n\n\n\n## Exponential Moving Averaging of weights(EMA)\n\nIt caluculates exponential moving average of weights during training. It is usually done on a minibatch level. I chose 10 updates per epoch for speed.\n[1], [2], [3] use EMA.\n\n\n## Tuning for ensemble\n\n### n_embed: #of pretrained embeddings in a single model.\nI tried n_embed=1, 2, 3, 4, and 2 is best. Using 4 embeddings is better for a single model F1, but lacks model diversity, which cause worse ensemble performance.\n\n### epoch\nThe followings are mean F1 of 10-fold CV. The best epoch is 5 for a single model and 8 for ensemble.\n![Image](https://raw.githubusercontent.com/tks0123456789/kaggle-Quora/master/cv10.png)\n\n\n[1] [A. Tarvainen and H. Valpola(2017) Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results.](https://arxiv.org/abs/1703.01780)\n\n[2] [Yu, A. W., Dohan, D., Luong, M.-T., Zhao, R., Chen, K., Norouzi, M., and Le, Q. V.(2018) QANet: Combining local convolution with global self-attention for reading comprehension](https://arxiv.org/abs/1804.09541)\n\n[3] [Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi (2016) Bidirectional attention flow for machine comprehension.](https://arxiv.org/abs/1611.01603)\n\n[4] [Douwe Kiela, Changhan Wang, Kyunghyun Cho (2018) Dynamic Meta-Embeddings for Improved Sentence Representations](https://arxiv.org/abs/1804.07983)",
    "472502": "Great solution! I mean \"great\" refers to difference from common shared kernel, like EMA, embedding combination variety for model ensembles. But I have several questions. Do you have an idea about how many improvements brought by PME and EMA? I have looked at your codes. I do not find the 10-fold cv part. It seems you do not do cross validation but how do you evaluate your model performance and reach the good balance between overfitting and underfitting? Do you publish the kernel that generates the 10-fold cv model performance plot against epoch?",
    "472765": "Congratulations @tks",
    "473361": "0.0015-0.002 for EMA. I don't know improvement by PME because I didn't try other embedding method. Code and result are [here](https://github.com/tks0123456789/kaggle-Quora)."
  },
  "source": "meta"
}