{
  "id": 71778,
  "title": "3 Methods to combine embeddings",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71778",
  "author_name": "Shujian Liu",
  "post_date": "2018-11-16T12:44:36.587000",
  "votes": 146,
  "comment_count": 29,
  "views": 0,
  "content": "<p>I would love to use some new context based embeddings like ELMO and BERT but it is not allowed. How can we take advantage of all the provided models? There are 3 major methods.</p>\n\n<ol>\n<li><p>Train the same model on separate embedding and blend. Great starter kernel by SRK:\n<a href=\"https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\">https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings</a></p></li>\n<li><p>Concat: I tried to follow Marios’ solution on Toxic competition by concatenating the embeddings:\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630</a>\nMy result is not as good as SKR’s kernel but I added a CNN model to beat the baseline: <a href=\"https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d\">https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d</a></p></li>\n<li><p>Meta embedding:\nThe better method is called meta embedding. There some cool papers from recent years:</p>\n\n<ul><li><p><a href=\"http://aclweb.org/anthology/D18-1176\">http://aclweb.org/anthology/D18-1176</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1808.04334.pdf\">https://arxiv.org/pdf/1808.04334.pdf</a></p></li>\n<li><p><a href=\"https://arxiv.org/abs/1508.04257\">https://arxiv.org/abs/1508.04257</a></p></li>\n<li><p><a href=\"http://aclweb.org/anthology/N18-2031\">http://aclweb.org/anthology/N18-2031</a></p></li>\n<li><p><a href=\"https://arxiv.org/abs/1704.01419\">https://arxiv.org/abs/1704.01419</a></p></li>\n<li><p><a href=\"http://aclweb.org/anthology/K18-1028\">http://aclweb.org/anthology/K18-1028</a></p></li></ul></li>\n</ol>\n\n<p>There are several methods but after some attempts, I decided to go with this paper: Frustratingly Easy Meta-Embedding – Computing Meta-Embeddings by Averaging Source Word Embeddings. Meta Embedding is a fancy name but It can be just a simple average of all embedding (I didn’t use word2vec since it is similar to FastText and it has low Val score).</p>\n\n<p>The conclusion from their paper: </p>\n\n<p>&gt; We have presented an argument for averaging as a valid meta-embedding technique, and found experimental performance to be close to, or in some cases better than that of concatenation, with the additional benefit of reduced dimensionality. We propose that when conducting meta-embedding, both concatenation and averaging should be considered as methods of combining embedding spaces, and their individual advantages considered.</p>\n\n<p>That sounds very efficient and useful since there is a time limit for the competition. We can put way more models in 2 hrs with this little trick. I am still working on this but here is some initial result my kernel: <a href=\"https://www.kaggle.com/shujian/mix-of-nn-models-based-on-meta-embedding\">https://www.kaggle.com/shujian/mix-of-nn-models-based-on-meta-embedding</a></p>",
  "messages": [
    {
      "id": 422577,
      "postDate": "2018-11-16T12:44:36.587Z",
      "content": "<p>I would love to use some new context based embeddings like ELMO and BERT but it is not allowed. How can we take advantage of all the provided models? There are 3 major methods.</p>\n\n<ol>\n<li><p>Train the same model on separate embedding and blend. Great starter kernel by SRK:\n<a href=\"https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\">https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings</a></p></li>\n<li><p>Concat: I tried to follow Marios’ solution on Toxic competition by concatenating the embeddings:\n<a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630</a>\nMy result is not as good as SKR’s kernel but I added a CNN model to beat the baseline: <a href=\"https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d\">https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d</a></p></li>\n<li><p>Meta embedding:\nThe better method is called meta embedding. There some cool papers from recent years:</p>\n\n<ul><li><p><a href=\"http://aclweb.org/anthology/D18-1176\">http://aclweb.org/anthology/D18-1176</a></p></li>\n<li><p><a href=\"https://arxiv.org/pdf/1808.04334.pdf\">https://arxiv.org/pdf/1808.04334.pdf</a></p></li>\n<li><p><a href=\"https://arxiv.org/abs/1508.04257\">https://arxiv.org/abs/1508.04257</a></p></li>\n<li><p><a href=\"http://aclweb.org/anthology/N18-2031\">http://aclweb.org/anthology/N18-2031</a></p></li>\n<li><p><a href=\"https://arxiv.org/abs/1704.01419\">https://arxiv.org/abs/1704.01419</a></p></li>\n<li><p><a href=\"http://aclweb.org/anthology/K18-1028\">http://aclweb.org/anthology/K18-1028</a></p></li></ul></li>\n</ol>\n\n<p>There are several methods but after some attempts, I decided to go with this paper: Frustratingly Easy Meta-Embedding – Computing Meta-Embeddings by Averaging Source Word Embeddings. Meta Embedding is a fancy name but It can be just a simple average of all embedding (I didn’t use word2vec since it is similar to FastText and it has low Val score).</p>\n\n<p>The conclusion from their paper: </p>\n\n<p>&gt; We have presented an argument for averaging as a valid meta-embedding technique, and found experimental performance to be close to, or in some cases better than that of concatenation, with the additional benefit of reduced dimensionality. We propose that when conducting meta-embedding, both concatenation and averaging should be considered as methods of combining embedding spaces, and their individual advantages considered.</p>\n\n<p>That sounds very efficient and useful since there is a time limit for the competition. We can put way more models in 2 hrs with this little trick. I am still working on this but here is some initial result my kernel: <a href=\"https://www.kaggle.com/shujian/mix-of-nn-models-based-on-meta-embedding\">https://www.kaggle.com/shujian/mix-of-nn-models-based-on-meta-embedding</a></p>",
      "rawMarkdown": "I would love to use some new context based embeddings like ELMO and BERT but it is not allowed. How can we take advantage of all the provided models? There are 3 major methods.\n\n1. Train the same model on separate embedding and blend. Great starter kernel by SRK:\nhttps://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\n\n2. Concat: I tried to follow Marios’ solution on Toxic competition by concatenating the embeddings:\nhttps://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630\nMy result is not as good as SKR’s kernel but I added a CNN model to beat the baseline: https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d\n\n\n3. Meta embedding:\nThe better method is called meta embedding. There some cool papers from recent years:\n\n - http://aclweb.org/anthology/D18-1176\n\n - https://arxiv.org/pdf/1808.04334.pdf\n\n - https://arxiv.org/abs/1508.04257\n\n - http://aclweb.org/anthology/N18-2031\n\n - https://arxiv.org/abs/1704.01419\n\n - http://aclweb.org/anthology/K18-1028\n\nThere are several methods but after some attempts, I decided to go with this paper: Frustratingly Easy Meta-Embedding – Computing Meta-Embeddings by Averaging Source Word Embeddings. Meta Embedding is a fancy name but It can be just a simple average of all embedding (I didn’t use word2vec since it is similar to FastText and it has low Val score).\n\nThe conclusion from their paper: \n\n&gt; We have presented an argument for averaging as a valid meta-embedding technique, and found experimental performance to be close to, or in some cases better than that of concatenation, with the additional benefit of reduced dimensionality. We propose that when conducting meta-embedding, both concatenation and averaging should be considered as methods of combining embedding spaces, and their individual advantages considered.\n\nThat sounds very efficient and useful since there is a time limit for the competition. We can put way more models in 2 hrs with this little trick. I am still working on this but here is some initial result my kernel: https://www.kaggle.com/shujian/mix-of-nn-models-based-on-meta-embedding",
      "votes": 145
    },
    {
      "id": 429078,
      "postDate": "2018-11-28T09:58:48.363Z",
      "content": "<p>If any one is interested in my keras implementation of the CDME from the first paper:</p>\n\n<pre><code>from keras.layers import Activation\nfrom keras.layers import multiply, Lambda\nimport keras.backend as K\n\ndef CDME_Block(inp, maxlen):\n    \"\"\"\n    # inp = tensor of shape (?,maxlen,embedding dim,n_emb)) n_emb is number of embedding matrices\n    # out = tensor of shape (?,maxlen,embedding dim)\n    \"\"\"\n    init = inp\n    x = Reshape((maxlen,-1))(inp)\n    x = CuDNNLSTM(n_emb,return_sequences = True)(x)\n    x = Activation('sigmoid')(x)\n    x = Reshape((maxlen,1,n_emb))(x)\n    x = multiply([init, x])\n    out = Lambda(lambda x: K.sum(x, axis=-1))(x)\n    return out\n</code></pre>\n\n<p>It basically merges several embedding matrices into one using word contect depended attention </p>",
      "rawMarkdown": "If any one is interested in my keras implementation of the CDME from the first paper:\n\n    from keras.layers import Activation\n    from keras.layers import multiply, Lambda\n    import keras.backend as K\n\n    def CDME_Block(inp, maxlen):\n        \"\"\"\n        # inp = tensor of shape (?,maxlen,embedding dim,n_emb)) n_emb is number of embedding matrices\n        # out = tensor of shape (?,maxlen,embedding dim)\n        \"\"\"\n        init = inp\n        x = Reshape((maxlen,-1))(inp)\n        x = CuDNNLSTM(n_emb,return_sequences = True)(x)\n        x = Activation('sigmoid')(x)\n        x = Reshape((maxlen,1,n_emb))(x)\n        x = multiply([init, x])\n        out = Lambda(lambda x: K.sum(x, axis=-1))(x)\n        return out\n\nIt basically merges several embedding matrices into one using word contect depended attention ",
      "votes": 19,
      "replies": [
        {
          "id": 429752,
          "postDate": "2018-11-29T09:54:09.473Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 430731,
          "postDate": "2018-11-30T22:32:30.567Z",
          "content": "<p>I am not sure how long it takes to merge, I used it to merge 4 embeddings and reduced training time needed by 10% compared to using concatenation, by having roughly the same accuracy </p>",
          "rawMarkdown": "I am not sure how long it takes to merge, I used it to merge 4 embeddings and reduced training time needed by 10% compared to using concatenation, by having roughly the same accuracy "
        },
        {
          "id": 597693,
          "postDate": "2019-08-12T17:04:03.463Z",
          "content": "<p>This is DME not CDME</p>",
          "rawMarkdown": "This is DME not CDME"
        }
      ]
    },
    {
      "id": 422587,
      "postDate": "2018-11-16T12:56:40.330Z",
      "content": "<p>So Insincere of you ;-) Great work and thanks for sharing valuable material !</p>",
      "rawMarkdown": "So Insincere of you ;-) Great work and thanks for sharing valuable material !",
      "votes": 3,
      "replies": [
        {
          "id": 422588,
          "postDate": "2018-11-16T13:00:24.983Z",
          "content": "<p>You are welcome!</p>",
          "rawMarkdown": "You are welcome!"
        }
      ]
    },
    {
      "id": 430045,
      "postDate": "2018-11-29T18:11:37.113Z",
      "content": "<p>Thank you so much for sharing this. I think I should start looking/implementing latest research more. Great learning:)</p>",
      "rawMarkdown": "Thank you so much for sharing this. I think I should start looking/implementing latest research more. Great learning:)",
      "votes": 1
    },
    {
      "id": 428514,
      "postDate": "2018-11-27T12:08:38.377Z",
      "content": "<p>awesome</p>",
      "rawMarkdown": "awesome",
      "votes": 1
    },
    {
      "id": 427755,
      "postDate": "2018-11-26T04:49:37.797Z",
      "content": "<p>Sounds great!</p>",
      "rawMarkdown": "Sounds great!",
      "votes": 1
    },
    {
      "id": 424536,
      "postDate": "2018-11-20T09:30:21.740Z",
      "content": "<p>Thanks for your insights and that collection of papers! I always think its funny and humbling when simple methods (like averaging) outperform or match more complex ones. </p>",
      "rawMarkdown": "Thanks for your insights and that collection of papers! I always think its funny and humbling when simple methods (like averaging) outperform or match more complex ones. ",
      "votes": 1
    },
    {
      "id": 424022,
      "postDate": "2018-11-19T12:55:27.280Z",
      "content": "<p>your kernels are really helpful !</p>",
      "rawMarkdown": "your kernels are really helpful !",
      "votes": 1
    },
    {
      "id": 423970,
      "postDate": "2018-11-19T10:53:05.383Z",
      "content": "<p>Thanks so much @Shujian Liu. Very useful papers.</p>",
      "rawMarkdown": "Thanks so much @Shujian Liu. Very useful papers.",
      "votes": 1
    },
    {
      "id": 423587,
      "postDate": "2018-11-18T16:21:07.860Z",
      "content": "<p>It sounds interesting :-)</p>",
      "rawMarkdown": "It sounds interesting :-)",
      "votes": 1
    },
    {
      "id": 423022,
      "postDate": "2018-11-17T09:36:57.800Z",
      "content": "<p>I have not read many papers. This is something I should do more often. I often come across these methods in stacking or blending like suggested here where averaging gives better results. What is the reason behind this? </p>",
      "rawMarkdown": "I have not read many papers. This is something I should do more often. I often come across these methods in stacking or blending like suggested here where averaging gives better results. What is the reason behind this? ",
      "votes": 1,
      "replies": [
        {
          "id": 423119,
          "postDate": "2018-11-17T14:36:22.317Z",
          "content": "<p>I guess stacking and blending can lead to overfiting sometime.</p>",
          "rawMarkdown": "I guess stacking and blending can lead to overfiting sometime.",
          "votes": 1
        }
      ]
    },
    {
      "id": 422622,
      "postDate": "2018-11-16T14:14:57.780Z",
      "content": "<p>Thanks for sharing !</p>\n\n<p>I did some research and I found this <a href=\"http://www.aaai.org/ocs/index.php/AAAI/AAAI16/paper/download/11777/11995\">http://www.aaai.org/ocs/index.php/AAAI/AAAI16/paper/download/11777/11995</a> (you've probably read it already)\nSome of their methods are interesting, I might give some of them a try on a basic model, to see which is better.</p>",
      "rawMarkdown": "Thanks for sharing !\n\nI did some research and I found this http://www.aaai.org/ocs/index.php/AAAI/AAAI16/paper/download/11777/11995 (you've probably read it already)\nSome of their methods are interesting, I might give some of them a try on a basic model, to see which is better.",
      "votes": 1,
      "replies": [
        {
          "id": 422626,
          "postDate": "2018-11-16T14:19:46.817Z",
          "content": "<p>Thanks for sharing. I haven't but I tried a paper from their team: Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation (best paper of conll 2018) (<a href=\"http://aclweb.org/anthology/K18-1028\">http://aclweb.org/anthology/K18-1028</a>) and didn't find it helpful here. </p>\n\n<p>Their code is simple (<a href=\"https://github.com/artetxem/uncovec/blob/master/post-process.py\">https://github.com/artetxem/uncovec/blob/master/post-process.py</a>):</p>\n\n<p>l, q = np.linalg.eigh(x.T.dot(x)); w = q*(l**args.alpha); x = x.dot(w)</p>",
          "rawMarkdown": "Thanks for sharing. I haven't but I tried a paper from their team: Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation (best paper of conll 2018) (http://aclweb.org/anthology/K18-1028) and didn't find it helpful here. \n\nTheir code is simple (https://github.com/artetxem/uncovec/blob/master/post-process.py):\n\nl, q = np.linalg.eigh(x.T.dot(x)); w = q*(l**args.alpha); x = x.dot(w)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 436128,
      "postDate": "2018-12-09T16:48:16.863Z",
      "content": "<p>Thanks. I was surprised that averaging the embeddings was giving good results because I couldn't understand what it meant to combine totally separate vector spaces with an average. The \"Frustratingly Easy Meta-Embedding\" paper helps make sense of that nonsense.</p>",
      "rawMarkdown": "Thanks. I was surprised that averaging the embeddings was giving good results because I couldn't understand what it meant to combine totally separate vector spaces with an average. The \"Frustratingly Easy Meta-Embedding\" paper helps make sense of that nonsense.",
      "votes": 2
    },
    {
      "id": 424619,
      "postDate": "2018-11-20T12:44:47.633Z",
      "content": "<p>You are so great no matter on this discussion or kernel. I have read all your kernels, you have tested with different models, different algorithms, different blending method. How you learn so quickly to make all this done? I admire you much!</p>",
      "rawMarkdown": "You are so great no matter on this discussion or kernel. I have read all your kernels, you have tested with different models, different algorithms, different blending method. How you learn so quickly to make all this done? I admire you much!",
      "votes": 2
    },
    {
      "id": 423251,
      "postDate": "2018-11-17T19:29:23.393Z",
      "content": "<p>Following your post i used the two approach : concatenate and Averaging and i had better result on test set with concatenate (using the 3 embeddings glove , wiki and param). Tried to use only glove and param when Averaging as you used but concatenate still better. My kernel is not  a great kernel (BI GRU + attention + CONV1D) just giving 0.679) but my two cents.</p>",
      "rawMarkdown": "Following your post i used the two approach : concatenate and Averaging and i had better result on test set with concatenate (using the 3 embeddings glove , wiki and param). Tried to use only glove and param when Averaging as you used but concatenate still better. My kernel is not  a great kernel (BI GRU + attention + CONV1D) just giving 0.679) but my two cents.",
      "votes": 2,
      "replies": [
        {
          "id": 423590,
          "postDate": "2018-11-18T16:29:42.503Z",
          "content": "<p>Yeah... concatenation has slightly better results but averaging is more efficient. That's what I was trying to do: put more models into 2 hrs :)</p>",
          "rawMarkdown": "Yeah... concatenation has slightly better results but averaging is more efficient. That's what I was trying to do: put more models into 2 hrs :)"
        }
      ]
    },
    {
      "id": 423209,
      "postDate": "2018-11-17T17:56:42.787Z",
      "content": "<p>Very Nice Sound great!</p>",
      "rawMarkdown": "Very Nice Sound great!",
      "votes": 1
    },
    {
      "id": 601757,
      "postDate": "2019-08-18T05:41:29.070Z",
      "content": "<p>thx for sharing</p>",
      "rawMarkdown": "thx for sharing"
    },
    {
      "id": 425339,
      "postDate": "2018-11-21T13:17:00.350Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    },
    {
      "id": 422984,
      "postDate": "2018-11-17T07:20:43.600Z",
      "content": "<p>Great Great Great, Thanks <a href=\"/shujian\">@shujian</a> </p>",
      "rawMarkdown": "Great Great Great, Thanks @shujian \n",
      "votes": 3
    },
    {
      "id": 430219,
      "postDate": "2018-11-30T03:16:21.087Z",
      "content": "<p>This is fabulous, Thanks <a href=\"/shujian\">@shujian</a></p>",
      "rawMarkdown": "This is fabulous, Thanks @shujian",
      "votes": 1
    },
    {
      "id": 428321,
      "postDate": "2018-11-27T04:38:46.593Z",
      "content": "<p>Awesome, thanks for sharing!</p>",
      "rawMarkdown": "Awesome, thanks for sharing!",
      "votes": 1
    },
    {
      "id": 426084,
      "postDate": "2018-11-22T15:22:42.450Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 423996,
      "postDate": "2018-11-19T11:57:38.850Z",
      "content": "<p>very informative. Thanks for sharing!</p>",
      "rawMarkdown": "very informative. Thanks for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 429078,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-11-28T09:58:48.363000",
      "content": "<p>If any one is interested in my keras implementation of the CDME from the first paper:</p>\n\n<pre><code>from keras.layers import Activation\nfrom keras.layers import multiply, Lambda\nimport keras.backend as K\n\ndef CDME_Block(inp, maxlen):\n    \"\"\"\n    # inp = tensor of shape (?,maxlen,embedding dim,n_emb)) n_emb is number of embedding matrices\n    # out = tensor of shape (?,maxlen,embedding dim)\n    \"\"\"\n    init = inp\n    x = Reshape((maxlen,-1))(inp)\n    x = CuDNNLSTM(n_emb,return_sequences = True)(x)\n    x = Activation('sigmoid')(x)\n    x = Reshape((maxlen,1,n_emb))(x)\n    x = multiply([init, x])\n    out = Lambda(lambda x: K.sum(x, axis=-1))(x)\n    return out\n</code></pre>\n\n<p>It basically merges several embedding matrices into one using word contect depended attention </p>",
      "votes": 19,
      "replies": [
        {
          "id": 429752,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-29T09:54:09.473000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430731,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-11-30T22:32:30.567000",
          "content": "<p>I am not sure how long it takes to merge, I used it to merge 4 embeddings and reduced training time needed by 10% compared to using concatenation, by having roughly the same accuracy </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 597693,
          "author_name": "gullit",
          "author_url": "",
          "post_date": "2019-08-12T17:04:03.463000",
          "content": "<p>This is DME not CDME</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422587,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-11-16T12:56:40.330000",
      "content": "<p>So Insincere of you ;-) Great work and thanks for sharing valuable material !</p>",
      "votes": 3,
      "replies": [
        {
          "id": 422588,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-16T13:00:24.983000",
          "content": "<p>You are welcome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 430045,
      "author_name": "jhajhria",
      "author_url": "",
      "post_date": "2018-11-29T18:11:37.113000",
      "content": "<p>Thank you so much for sharing this. I think I should start looking/implementing latest research more. Great learning:)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 428514,
      "author_name": "Victor An",
      "author_url": "",
      "post_date": "2018-11-27T12:08:38.377000",
      "content": "<p>awesome</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 427755,
      "author_name": "PrakharMishra",
      "author_url": "",
      "post_date": "2018-11-26T04:49:37.797000",
      "content": "<p>Sounds great!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 424536,
      "author_name": "Marlon Flügge",
      "author_url": "",
      "post_date": "2018-11-20T09:30:21.740000",
      "content": "<p>Thanks for your insights and that collection of papers! I always think its funny and humbling when simple methods (like averaging) outperform or match more complex ones. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 424022,
      "author_name": "Kartikeya Bhardwaj",
      "author_url": "",
      "post_date": "2018-11-19T12:55:27.280000",
      "content": "<p>your kernels are really helpful !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 423970,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-11-19T10:53:05.383000",
      "content": "<p>Thanks so much @Shujian Liu. Very useful papers.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 423587,
      "author_name": "Kenshiro75",
      "author_url": "",
      "post_date": "2018-11-18T16:21:07.860000",
      "content": "<p>It sounds interesting :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 423022,
      "author_name": "Siddharth Yadav",
      "author_url": "",
      "post_date": "2018-11-17T09:36:57.800000",
      "content": "<p>I have not read many papers. This is something I should do more often. I often come across these methods in stacking or blending like suggested here where averaging gives better results. What is the reason behind this? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 423119,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-17T14:36:22.317000",
          "content": "<p>I guess stacking and blending can lead to overfiting sometime.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 422622,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2018-11-16T14:14:57.780000",
      "content": "<p>Thanks for sharing !</p>\n\n<p>I did some research and I found this <a href=\"http://www.aaai.org/ocs/index.php/AAAI/AAAI16/paper/download/11777/11995\">http://www.aaai.org/ocs/index.php/AAAI/AAAI16/paper/download/11777/11995</a> (you've probably read it already)\nSome of their methods are interesting, I might give some of them a try on a basic model, to see which is better.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 422626,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-16T14:19:46.817000",
          "content": "<p>Thanks for sharing. I haven't but I tried a paper from their team: Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation (best paper of conll 2018) (<a href=\"http://aclweb.org/anthology/K18-1028\">http://aclweb.org/anthology/K18-1028</a>) and didn't find it helpful here. </p>\n\n<p>Their code is simple (<a href=\"https://github.com/artetxem/uncovec/blob/master/post-process.py\">https://github.com/artetxem/uncovec/blob/master/post-process.py</a>):</p>\n\n<p>l, q = np.linalg.eigh(x.T.dot(x)); w = q*(l**args.alpha); x = x.dot(w)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 436128,
      "author_name": "HappyPrancer",
      "author_url": "",
      "post_date": "2018-12-09T16:48:16.863000",
      "content": "<p>Thanks. I was surprised that averaging the embeddings was giving good results because I couldn't understand what it meant to combine totally separate vector spaces with an average. The \"Frustratingly Easy Meta-Embedding\" paper helps make sense of that nonsense.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 424619,
      "author_name": "Guangqiang Lu",
      "author_url": "",
      "post_date": "2018-11-20T12:44:47.633000",
      "content": "<p>You are so great no matter on this discussion or kernel. I have read all your kernels, you have tested with different models, different algorithms, different blending method. How you learn so quickly to make all this done? I admire you much!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 423251,
      "author_name": "Jean-Marc",
      "author_url": "",
      "post_date": "2018-11-17T19:29:23.393000",
      "content": "<p>Following your post i used the two approach : concatenate and Averaging and i had better result on test set with concatenate (using the 3 embeddings glove , wiki and param). Tried to use only glove and param when Averaging as you used but concatenate still better. My kernel is not  a great kernel (BI GRU + attention + CONV1D) just giving 0.679) but my two cents.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 423590,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2018-11-18T16:29:42.503000",
          "content": "<p>Yeah... concatenation has slightly better results but averaging is more efficient. That's what I was trying to do: put more models into 2 hrs :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 423209,
      "author_name": "Arunkumar Venkataramanan",
      "author_url": "",
      "post_date": "2018-11-17T17:56:42.787000",
      "content": "<p>Very Nice Sound great!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 601757,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-08-18T05:41:29.070000",
      "content": "<p>thx for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 425339,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-21T13:17:00.350000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 422984,
      "author_name": "MJ Bahmani",
      "author_url": "",
      "post_date": "2018-11-17T07:20:43.600000",
      "content": "<p>Great Great Great, Thanks <a href=\"/shujian\">@shujian</a> </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 430219,
      "author_name": "Vishy",
      "author_url": "",
      "post_date": "2018-11-30T03:16:21.087000",
      "content": "<p>This is fabulous, Thanks <a href=\"/shujian\">@shujian</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 428321,
      "author_name": "Hanc",
      "author_url": "",
      "post_date": "2018-11-27T04:38:46.593000",
      "content": "<p>Awesome, thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 426084,
      "author_name": "Palasanu George",
      "author_url": "",
      "post_date": "2018-11-22T15:22:42.450000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 423996,
      "author_name": "Mayank",
      "author_url": "",
      "post_date": "2018-11-19T11:57:38.850000",
      "content": "<p>very informative. Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422577": "I would love to use some new context based embeddings like ELMO and BERT but it is not allowed. How can we take advantage of all the provided models? There are 3 major methods.\n\n1. Train the same model on separate embedding and blend. Great starter kernel by SRK:\nhttps://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings\n\n2. Concat: I tried to follow Marios’ solution on Toxic competition by concatenating the embeddings:\nhttps://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630\nMy result is not as good as SKR’s kernel but I added a CNN model to beat the baseline: https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d\n\n\n3. Meta embedding:\nThe better method is called meta embedding. There some cool papers from recent years:\n\n - http://aclweb.org/anthology/D18-1176\n\n - https://arxiv.org/pdf/1808.04334.pdf\n\n - https://arxiv.org/abs/1508.04257\n\n - http://aclweb.org/anthology/N18-2031\n\n - https://arxiv.org/abs/1704.01419\n\n - http://aclweb.org/anthology/K18-1028\n\nThere are several methods but after some attempts, I decided to go with this paper: Frustratingly Easy Meta-Embedding – Computing Meta-Embeddings by Averaging Source Word Embeddings. Meta Embedding is a fancy name but It can be just a simple average of all embedding (I didn’t use word2vec since it is similar to FastText and it has low Val score).\n\nThe conclusion from their paper: \n\n&gt; We have presented an argument for averaging as a valid meta-embedding technique, and found experimental performance to be close to, or in some cases better than that of concatenation, with the additional benefit of reduced dimensionality. We propose that when conducting meta-embedding, both concatenation and averaging should be considered as methods of combining embedding spaces, and their individual advantages considered.\n\nThat sounds very efficient and useful since there is a time limit for the competition. We can put way more models in 2 hrs with this little trick. I am still working on this but here is some initial result my kernel: https://www.kaggle.com/shujian/mix-of-nn-models-based-on-meta-embedding",
    "429078": "If any one is interested in my keras implementation of the CDME from the first paper:\n\n    from keras.layers import Activation\n    from keras.layers import multiply, Lambda\n    import keras.backend as K\n\n    def CDME_Block(inp, maxlen):\n        \"\"\"\n        # inp = tensor of shape (?,maxlen,embedding dim,n_emb)) n_emb is number of embedding matrices\n        # out = tensor of shape (?,maxlen,embedding dim)\n        \"\"\"\n        init = inp\n        x = Reshape((maxlen,-1))(inp)\n        x = CuDNNLSTM(n_emb,return_sequences = True)(x)\n        x = Activation('sigmoid')(x)\n        x = Reshape((maxlen,1,n_emb))(x)\n        x = multiply([init, x])\n        out = Lambda(lambda x: K.sum(x, axis=-1))(x)\n        return out\n\nIt basically merges several embedding matrices into one using word contect depended attention ",
    "422587": "So Insincere of you ;-) Great work and thanks for sharing valuable material !",
    "430045": "Thank you so much for sharing this. I think I should start looking/implementing latest research more. Great learning:)",
    "428514": "awesome",
    "427755": "Sounds great!",
    "424536": "Thanks for your insights and that collection of papers! I always think its funny and humbling when simple methods (like averaging) outperform or match more complex ones. ",
    "424022": "your kernels are really helpful !",
    "423970": "Thanks so much @Shujian Liu. Very useful papers.",
    "423587": "It sounds interesting :-)",
    "423022": "I have not read many papers. This is something I should do more often. I often come across these methods in stacking or blending like suggested here where averaging gives better results. What is the reason behind this? ",
    "422622": "Thanks for sharing !\n\nI did some research and I found this http://www.aaai.org/ocs/index.php/AAAI/AAAI16/paper/download/11777/11995 (you've probably read it already)\nSome of their methods are interesting, I might give some of them a try on a basic model, to see which is better.",
    "436128": "Thanks. I was surprised that averaging the embeddings was giving good results because I couldn't understand what it meant to combine totally separate vector spaces with an average. The \"Frustratingly Easy Meta-Embedding\" paper helps make sense of that nonsense.",
    "424619": "You are so great no matter on this discussion or kernel. I have read all your kernels, you have tested with different models, different algorithms, different blending method. How you learn so quickly to make all this done? I admire you much!",
    "423251": "Following your post i used the two approach : concatenate and Averaging and i had better result on test set with concatenate (using the 3 embeddings glove , wiki and param). Tried to use only glove and param when Averaging as you used but concatenate still better. My kernel is not  a great kernel (BI GRU + attention + CONV1D) just giving 0.679) but my two cents.",
    "423209": "Very Nice Sound great!",
    "601757": "thx for sharing",
    "425339": "",
    "422984": "Great Great Great, Thanks @shujian \n",
    "430219": "This is fabulous, Thanks @shujian",
    "428321": "Awesome, thanks for sharing!",
    "426084": "Thanks for sharing!",
    "423996": "very informative. Thanks for sharing!"
  }
}