{
  "id": 78533,
  "title": "Implementing Research Paper",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78533",
  "author_name": "",
  "post_date": "2019-01-25T04:42:30.406450100Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,\n      I am trying to implement the paper ' <a href=\"https://link.springer.com/article/10.1007/s41060-017-0088-4\">https://link.springer.com/article/10.1007/s41060-017-0088-4</a> '. Since we don't have DSSM word to vector model, I used Glove word embedding. I preprocessed the input data and then created the model, which i believe matches with the architecture mentioned in the paper. But when i implemented it gives the  f1 score as 0.11 as mentioned below. Can someone please tell me where am i going wrong and how to correct it ? </p>\n\n<p>OUTPUT:</p>\n\n<pre><code>_________________________________________________________________\nLayer (type)                 Output Shape              Param #   \n=================================================================\ninput_1 (InputLayer)         (None, 70)                0         \n_________________________________________________________________\nembedding_1 (Embedding)      (None, 70, 300)           73146900  \n_________________________________________________________________\nreshape_1 (Reshape)          (None, 70, 300, 1)        0         \n_________________________________________________________________\nconv2d_1 (Conv2D)            (None, 68, 276, 4)        304       \n_________________________________________________________________\nconv2d_2 (Conv2D)            (None, 66, 252, 2)        602       \n_________________________________________________________________\nconv2d_3 (Conv2D)            (None, 64, 228, 1)        151       \n_________________________________________________________________\nreshape_2 (Reshape)          (None, 64, 228)           0         \n_________________________________________________________________\nbidirectional_1 (Bidirection (None, 32)                31488     \n_________________________________________________________________\ndense_1 (Dense)              (None, 4)                 132       \n_________________________________________________________________\ndense_2 (Dense)              (None, 1)                 5         \n=================================================================\nTotal params: 73,179,582\nTrainable params: 32,682\nNon-trainable params: 73,146,900\n_________________________________________________________________\nNone\nTrain on 1110204 samples, validate on 195918 samples\nEpoch 1/4\n - 1424s - loss: 14.9560 - val_loss: 14.9561\n\nEpoch 00001: val_loss improved from inf to 14.95606, saving model to weights_best_glove1.hdf5\nEpoch 2/4\n - 1413s - loss: 14.9560 - val_loss: 14.9561\n\nEpoch 00002: val_loss did not improve from 14.95606\n\nEpoch 00002: ReduceLROnPlateau reducing learning rate to 0.030000000447034835.\nEpoch 3/4\n - 1413s - loss: 14.9560 - val_loss: 14.9561\n\nEpoch 00003: val_loss did not improve from 14.95606\n\nEpoch 00003: ReduceLROnPlateau reducing learning rate to 0.018000000715255735.\nEpoch 00003: early stopping\noptimal F1: 0.1165 at threshold: 0.1000\n</code></pre>",
  "messages": [
    {
      "id": "461047",
      "postDate": "01/25/2019 04:42:30",
      "content": "<p>Hi,\n      I am trying to implement the paper ' <a href=\"https://link.springer.com/article/10.1007/s41060-017-0088-4\">https://link.springer.com/article/10.1007/s41060-017-0088-4</a> '. Since we don't have DSSM word to vector model, I used Glove word embedding. I preprocessed the input data and then created the model, which i believe matches with the architecture mentioned in the paper. But when i implemented it gives the  f1 score as 0.11 as mentioned below. Can someone please tell me where am i going wrong and how to correct it ? </p>\n\n<p>OUTPUT:</p>\n\n<pre><code>_________________________________________________________________\nLayer (type)                 Output Shape              Param #   \n=================================================================\ninput_1 (InputLayer)         (None, 70)                0         \n_________________________________________________________________\nembedding_1 (Embedding)      (None, 70, 300)           73146900  \n_________________________________________________________________\nreshape_1 (Reshape)          (None, 70, 300, 1)        0         \n_________________________________________________________________\nconv2d_1 (Conv2D)            (None, 68, 276, 4)        304       \n_________________________________________________________________\nconv2d_2 (Conv2D)            (None, 66, 252, 2)        602       \n_________________________________________________________________\nconv2d_3 (Conv2D)            (None, 64, 228, 1)        151       \n_________________________________________________________________\nreshape_2 (Reshape)          (None, 64, 228)           0         \n_________________________________________________________________\nbidirectional_1 (Bidirection (None, 32)                31488     \n_________________________________________________________________\ndense_1 (Dense)              (None, 4)                 132       \n_________________________________________________________________\ndense_2 (Dense)              (None, 1)                 5         \n=================================================================\nTotal params: 73,179,582\nTrainable params: 32,682\nNon-trainable params: 73,146,900\n_________________________________________________________________\nNone\nTrain on 1110204 samples, validate on 195918 samples\nEpoch 1/4\n - 1424s - loss: 14.9560 - val_loss: 14.9561\n\nEpoch 00001: val_loss improved from inf to 14.95606, saving model to weights_best_glove1.hdf5\nEpoch 2/4\n - 1413s - loss: 14.9560 - val_loss: 14.9561\n\nEpoch 00002: val_loss did not improve from 14.95606\n\nEpoch 00002: ReduceLROnPlateau reducing learning rate to 0.030000000447034835.\nEpoch 3/4\n - 1413s - loss: 14.9560 - val_loss: 14.9561\n\nEpoch 00003: val_loss did not improve from 14.95606\n\nEpoch 00003: ReduceLROnPlateau reducing learning rate to 0.018000000715255735.\nEpoch 00003: early stopping\noptimal F1: 0.1165 at threshold: 0.1000\n</code></pre>",
      "rawMarkdown": "Hi,\n      I am trying to implement the paper ' https://link.springer.com/article/10.1007/s41060-017-0088-4 '. Since we don't have DSSM word to vector model, I used Glove word embedding. I preprocessed the input data and then created the model, which i believe matches with the architecture mentioned in the paper. But when i implemented it gives the  f1 score as 0.11 as mentioned below. Can someone please tell me where am i going wrong and how to correct it ? \n\n\nOUTPUT:\n\n    _________________________________________________________________\n    Layer (type)                 Output Shape              Param #   \n    =================================================================\n    input_1 (InputLayer)         (None, 70)                0         \n    _________________________________________________________________\n    embedding_1 (Embedding)      (None, 70, 300)           73146900  \n    _________________________________________________________________\n    reshape_1 (Reshape)          (None, 70, 300, 1)        0         \n    _________________________________________________________________\n    conv2d_1 (Conv2D)            (None, 68, 276, 4)        304       \n    _________________________________________________________________\n    conv2d_2 (Conv2D)            (None, 66, 252, 2)        602       \n    _________________________________________________________________\n    conv2d_3 (Conv2D)            (None, 64, 228, 1)        151       \n    _________________________________________________________________\n    reshape_2 (Reshape)          (None, 64, 228)           0         \n    _________________________________________________________________\n    bidirectional_1 (Bidirection (None, 32)                31488     \n    _________________________________________________________________\n    dense_1 (Dense)              (None, 4)                 132       \n    _________________________________________________________________\n    dense_2 (Dense)              (None, 1)                 5         \n    =================================================================\n    Total params: 73,179,582\n    Trainable params: 32,682\n    Non-trainable params: 73,146,900\n    _________________________________________________________________\n    None\n    Train on 1110204 samples, validate on 195918 samples\n    Epoch 1/4\n     - 1424s - loss: 14.9560 - val_loss: 14.9561\n    \n    Epoch 00001: val_loss improved from inf to 14.95606, saving model to weights_best_glove1.hdf5\n    Epoch 2/4\n     - 1413s - loss: 14.9560 - val_loss: 14.9561\n    \n    Epoch 00002: val_loss did not improve from 14.95606\n    \n    Epoch 00002: ReduceLROnPlateau reducing learning rate to 0.030000000447034835.\n    Epoch 3/4\n     - 1413s - loss: 14.9560 - val_loss: 14.9561\n    \n    Epoch 00003: val_loss did not improve from 14.95606\n    \n    Epoch 00003: ReduceLROnPlateau reducing learning rate to 0.018000000715255735.\n    Epoch 00003: early stopping\n    optimal F1: 0.1165 at threshold: 0.1000",
      "votes": null
    },
    {
      "id": "461387",
      "postDate": "01/25/2019 22:21:22",
      "content": "<p>May be because they embedding they are using is a character level on.</p>\n\n<p>So the sentence, \"what is the weather is UK in December?\" is broken down as following and have a meaning over the time-distribution:</p>\n\n<p>wha</p>\n\n<p>hat</p>\n\n<p>at </p>\n\n<p>t i</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>mber</p>\n\n<p>ber</p>\n\n<p>er? </p>\n\n<p>On other hand, we are using complete words as embedding, so the con2d kernels can't correlate between input as output.</p>",
      "rawMarkdown": "May be because they embedding they are using is a character level on.\n\nSo the sentence, \"what is the weather is UK in December?\" is broken down as following and have a meaning over the time-distribution:\n\nwha\n\nhat\n\nat \n\nt i\n\n.\n\n.\n\n.\n\nmber\n\nber\n\ner? \n\n\nOn other hand, we are using complete words as embedding, so the con2d kernels can't correlate between input as output.",
      "votes": null
    },
    {
      "id": "461576",
      "postDate": "01/26/2019 12:18:27",
      "content": "<p>It's not learning anything, the loss stays the same. It's likely there is something wrong with your implementation. Can you post the code of how you create your model?</p>",
      "rawMarkdown": "It's not learning anything, the loss stays the same. It's likely there is something wrong with your implementation. Can you post the code of how you create your model?",
      "votes": null
    },
    {
      "id": "461676",
      "postDate": "01/26/2019 17:44:50",
      "content": "<p>Sure. Here it is.. \nHere max_words is the max no of words in a sentence which is initialized to 70. </p>\n\n<pre><code>def get_model(embedding_matrix):\n\n    print(embedding_matrix.shape);\n    num_filters = [4, 2, 1];\n\n    inp = kl.Input(shape=(max_words,))\n\n    x = kl.Embedding(input_dim = cl+1, output_dim = EMBEDDING_DIM, weights=[embedding_matrix], trainable = False)(inp)\n    x = kl.Reshape((max_words, EMBEDDING_DIM, 1))(x)\n\n    filter_size = (3, 25);  final_size = (64, 228); att_length = final_size[0]; inp_shape = (64, 228, 1);\n\n    x = kl.Conv2D(num_filters[0], kernel_size = filter_size, activation='relu')(x)\n    x = kl.Conv2D(num_filters[1], kernel_size = filter_size, activation='relu')(x)\n    x = kl.Conv2D(num_filters[2], kernel_size = filter_size, activation='relu')(x)\n\n    x = kl.Reshape(final_size, input_shape = inp_shape)(x);\n    print(x.get_shape());\n    x = kl.Bidirectional(kl.CuDNNLSTM(16, kernel_initializer=glorot_normal(seed=1029), recurrent_initializer=orthogonal(gain=1.0, seed=1029), name = 'bi_lstm1'))(x);\n\n    x = kl.Dense(4, activation=\"relu\")(x);\n\n    x = kl.Dense(1, activation = 'softmax')(x);\n    optim = keras.optimizers.Adagrad(lr=0.05, epsilon=1e-8, decay=0.0)\n\n    model = Model(inputs=inp, outputs=x)\n    model.compile(loss = 'binary_crossentropy', optimizer = optim);\n\n    print(model.summary());\n\n    return model;\n</code></pre>",
      "rawMarkdown": "Sure. Here it is.. \nHere max_words is the max no of words in a sentence which is initialized to 70. \n\n\n    def get_model(embedding_matrix):\n        \n        print(embedding_matrix.shape);\n        num_filters = [4, 2, 1];\n    \n        inp = kl.Input(shape=(max_words,))\n        \n        x = kl.Embedding(input_dim = cl+1, output_dim = EMBEDDING_DIM, weights=[embedding_matrix], trainable = False)(inp)\n        x = kl.Reshape((max_words, EMBEDDING_DIM, 1))(x)\n        \n        filter_size = (3, 25);  final_size = (64, 228); att_length = final_size[0]; inp_shape = (64, 228, 1);\n        \n        x = kl.Conv2D(num_filters[0], kernel_size = filter_size, activation='relu')(x)\n        x = kl.Conv2D(num_filters[1], kernel_size = filter_size, activation='relu')(x)\n        x = kl.Conv2D(num_filters[2], kernel_size = filter_size, activation='relu')(x)\n        \n        x = kl.Reshape(final_size, input_shape = inp_shape)(x);\n        print(x.get_shape());\n        x = kl.Bidirectional(kl.CuDNNLSTM(16, kernel_initializer=glorot_normal(seed=1029), recurrent_initializer=orthogonal(gain=1.0, seed=1029), name = 'bi_lstm1'))(x);\n        \n        x = kl.Dense(4, activation=\"relu\")(x);\n        \n        x = kl.Dense(1, activation = 'softmax')(x);\n        optim = keras.optimizers.Adagrad(lr=0.05, epsilon=1e-8, decay=0.0)\n        \n        model = Model(inputs=inp, outputs=x)\n        model.compile(loss = 'binary_crossentropy', optimizer = optim);\n        \n        print(model.summary());\n        \n        return model;",
      "votes": null
    },
    {
      "id": "461680",
      "postDate": "01/26/2019 17:50:43",
      "content": "<p>But even though , a word vector is the representation of the word in both embeddings. When they are separating 25 dimensional filter instead of 300 dimensions( which completely represents a word), if it can understand DSSM, even though it is character level trained, cant it correlate with glove ? Can you plz elaborate on this ?</p>",
      "rawMarkdown": "But even though , a word vector is the representation of the word in both embeddings. When they are separating 25 dimensional filter instead of 300 dimensions( which completely represents a word), if it can understand DSSM, even though it is character level trained, cant it correlate with glove ? Can you plz elaborate on this ?",
      "votes": null
    },
    {
      "id": "461746",
      "postDate": "01/26/2019 22:56:24",
      "content": "<p>I gave the model in the paper a try, but it can't learn anything even at lower learning rates. If u apply the lr=0.05 as described in the paper, the loss overshoots. Using a lower lr makes the loss converge but the model doesn't learn anything of value.</p>\n\n<p>Read up on DSSM if you can, but from my understanding, the DSSM embedding conserves a temporal relational between the adjacent steps of the embedding layer outcome and hence the model is able to learn. </p>\n\n<p>You can also try using Cov1D to get similar outcome, however in my experience it takes a long time to train it that way and wouldn't be feasible in this challenge.</p>",
      "rawMarkdown": "I gave the model in the paper a try, but it can't learn anything even at lower learning rates. If u apply the lr=0.05 as described in the paper, the loss overshoots. Using a lower lr makes the loss converge but the model doesn't learn anything of value.\n\nRead up on DSSM if you can, but from my understanding, the DSSM embedding conserves a temporal relational between the adjacent steps of the embedding layer outcome and hence the model is able to learn. \n\nYou can also try using Cov1D to get similar outcome, however in my experience it takes a long time to train it that way and wouldn't be feasible in this challenge.",
      "votes": null
    },
    {
      "id": "461860",
      "postDate": "01/27/2019 07:18:28",
      "content": "<p>Oh! Ok. I will read about it </p>",
      "rawMarkdown": "Oh! Ok. I will read about it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 461387,
      "author_name": "abdurrafae",
      "author_url": "",
      "post_date": "01/25/2019 22:21:22",
      "content": "<p>May be because they embedding they are using is a character level on.</p>\n\n<p>So the sentence, \"what is the weather is UK in December?\" is broken down as following and have a meaning over the time-distribution:</p>\n\n<p>wha</p>\n\n<p>hat</p>\n\n<p>at </p>\n\n<p>t i</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>.</p>\n\n<p>mber</p>\n\n<p>ber</p>\n\n<p>er? </p>\n\n<p>On other hand, we are using complete words as embedding, so the con2d kernels can't correlate between input as output.</p>",
      "votes": null,
      "replies": [
        {
          "id": 461680,
          "author_name": "ajayhayagreeve",
          "author_url": "",
          "post_date": "01/26/2019 17:50:43",
          "content": "<p>But even though , a word vector is the representation of the word in both embeddings. When they are separating 25 dimensional filter instead of 300 dimensions( which completely represents a word), if it can understand DSSM, even though it is character level trained, cant it correlate with glove ? Can you plz elaborate on this ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461746,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "01/26/2019 22:56:24",
          "content": "<p>I gave the model in the paper a try, but it can't learn anything even at lower learning rates. If u apply the lr=0.05 as described in the paper, the loss overshoots. Using a lower lr makes the loss converge but the model doesn't learn anything of value.</p>\n\n<p>Read up on DSSM if you can, but from my understanding, the DSSM embedding conserves a temporal relational between the adjacent steps of the embedding layer outcome and hence the model is able to learn. </p>\n\n<p>You can also try using Cov1D to get similar outcome, however in my experience it takes a long time to train it that way and wouldn't be feasible in this challenge.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461860,
          "author_name": "ajayhayagreeve",
          "author_url": "",
          "post_date": "01/27/2019 07:18:28",
          "content": "<p>Oh! Ok. I will read about it </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 461576,
      "author_name": "mschumacher",
      "author_url": "",
      "post_date": "01/26/2019 12:18:27",
      "content": "<p>It's not learning anything, the loss stays the same. It's likely there is something wrong with your implementation. Can you post the code of how you create your model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 461676,
          "author_name": "ajayhayagreeve",
          "author_url": "",
          "post_date": "01/26/2019 17:44:50",
          "content": "<p>Sure. Here it is.. \nHere max_words is the max no of words in a sentence which is initialized to 70. </p>\n\n<pre><code>def get_model(embedding_matrix):\n\n    print(embedding_matrix.shape);\n    num_filters = [4, 2, 1];\n\n    inp = kl.Input(shape=(max_words,))\n\n    x = kl.Embedding(input_dim = cl+1, output_dim = EMBEDDING_DIM, weights=[embedding_matrix], trainable = False)(inp)\n    x = kl.Reshape((max_words, EMBEDDING_DIM, 1))(x)\n\n    filter_size = (3, 25);  final_size = (64, 228); att_length = final_size[0]; inp_shape = (64, 228, 1);\n\n    x = kl.Conv2D(num_filters[0], kernel_size = filter_size, activation='relu')(x)\n    x = kl.Conv2D(num_filters[1], kernel_size = filter_size, activation='relu')(x)\n    x = kl.Conv2D(num_filters[2], kernel_size = filter_size, activation='relu')(x)\n\n    x = kl.Reshape(final_size, input_shape = inp_shape)(x);\n    print(x.get_shape());\n    x = kl.Bidirectional(kl.CuDNNLSTM(16, kernel_initializer=glorot_normal(seed=1029), recurrent_initializer=orthogonal(gain=1.0, seed=1029), name = 'bi_lstm1'))(x);\n\n    x = kl.Dense(4, activation=\"relu\")(x);\n\n    x = kl.Dense(1, activation = 'softmax')(x);\n    optim = keras.optimizers.Adagrad(lr=0.05, epsilon=1e-8, decay=0.0)\n\n    model = Model(inputs=inp, outputs=x)\n    model.compile(loss = 'binary_crossentropy', optimizer = optim);\n\n    print(model.summary());\n\n    return model;\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "461047": "Hi,\n      I am trying to implement the paper ' https://link.springer.com/article/10.1007/s41060-017-0088-4 '. Since we don't have DSSM word to vector model, I used Glove word embedding. I preprocessed the input data and then created the model, which i believe matches with the architecture mentioned in the paper. But when i implemented it gives the  f1 score as 0.11 as mentioned below. Can someone please tell me where am i going wrong and how to correct it ? \n\n\nOUTPUT:\n\n    _________________________________________________________________\n    Layer (type)                 Output Shape              Param #   \n    =================================================================\n    input_1 (InputLayer)         (None, 70)                0         \n    _________________________________________________________________\n    embedding_1 (Embedding)      (None, 70, 300)           73146900  \n    _________________________________________________________________\n    reshape_1 (Reshape)          (None, 70, 300, 1)        0         \n    _________________________________________________________________\n    conv2d_1 (Conv2D)            (None, 68, 276, 4)        304       \n    _________________________________________________________________\n    conv2d_2 (Conv2D)            (None, 66, 252, 2)        602       \n    _________________________________________________________________\n    conv2d_3 (Conv2D)            (None, 64, 228, 1)        151       \n    _________________________________________________________________\n    reshape_2 (Reshape)          (None, 64, 228)           0         \n    _________________________________________________________________\n    bidirectional_1 (Bidirection (None, 32)                31488     \n    _________________________________________________________________\n    dense_1 (Dense)              (None, 4)                 132       \n    _________________________________________________________________\n    dense_2 (Dense)              (None, 1)                 5         \n    =================================================================\n    Total params: 73,179,582\n    Trainable params: 32,682\n    Non-trainable params: 73,146,900\n    _________________________________________________________________\n    None\n    Train on 1110204 samples, validate on 195918 samples\n    Epoch 1/4\n     - 1424s - loss: 14.9560 - val_loss: 14.9561\n    \n    Epoch 00001: val_loss improved from inf to 14.95606, saving model to weights_best_glove1.hdf5\n    Epoch 2/4\n     - 1413s - loss: 14.9560 - val_loss: 14.9561\n    \n    Epoch 00002: val_loss did not improve from 14.95606\n    \n    Epoch 00002: ReduceLROnPlateau reducing learning rate to 0.030000000447034835.\n    Epoch 3/4\n     - 1413s - loss: 14.9560 - val_loss: 14.9561\n    \n    Epoch 00003: val_loss did not improve from 14.95606\n    \n    Epoch 00003: ReduceLROnPlateau reducing learning rate to 0.018000000715255735.\n    Epoch 00003: early stopping\n    optimal F1: 0.1165 at threshold: 0.1000",
    "461387": "May be because they embedding they are using is a character level on.\n\nSo the sentence, \"what is the weather is UK in December?\" is broken down as following and have a meaning over the time-distribution:\n\nwha\n\nhat\n\nat \n\nt i\n\n.\n\n.\n\n.\n\nmber\n\nber\n\ner? \n\n\nOn other hand, we are using complete words as embedding, so the con2d kernels can't correlate between input as output.",
    "461576": "It's not learning anything, the loss stays the same. It's likely there is something wrong with your implementation. Can you post the code of how you create your model?",
    "461676": "Sure. Here it is.. \nHere max_words is the max no of words in a sentence which is initialized to 70. \n\n\n    def get_model(embedding_matrix):\n        \n        print(embedding_matrix.shape);\n        num_filters = [4, 2, 1];\n    \n        inp = kl.Input(shape=(max_words,))\n        \n        x = kl.Embedding(input_dim = cl+1, output_dim = EMBEDDING_DIM, weights=[embedding_matrix], trainable = False)(inp)\n        x = kl.Reshape((max_words, EMBEDDING_DIM, 1))(x)\n        \n        filter_size = (3, 25);  final_size = (64, 228); att_length = final_size[0]; inp_shape = (64, 228, 1);\n        \n        x = kl.Conv2D(num_filters[0], kernel_size = filter_size, activation='relu')(x)\n        x = kl.Conv2D(num_filters[1], kernel_size = filter_size, activation='relu')(x)\n        x = kl.Conv2D(num_filters[2], kernel_size = filter_size, activation='relu')(x)\n        \n        x = kl.Reshape(final_size, input_shape = inp_shape)(x);\n        print(x.get_shape());\n        x = kl.Bidirectional(kl.CuDNNLSTM(16, kernel_initializer=glorot_normal(seed=1029), recurrent_initializer=orthogonal(gain=1.0, seed=1029), name = 'bi_lstm1'))(x);\n        \n        x = kl.Dense(4, activation=\"relu\")(x);\n        \n        x = kl.Dense(1, activation = 'softmax')(x);\n        optim = keras.optimizers.Adagrad(lr=0.05, epsilon=1e-8, decay=0.0)\n        \n        model = Model(inputs=inp, outputs=x)\n        model.compile(loss = 'binary_crossentropy', optimizer = optim);\n        \n        print(model.summary());\n        \n        return model;",
    "461680": "But even though , a word vector is the representation of the word in both embeddings. When they are separating 25 dimensional filter instead of 300 dimensions( which completely represents a word), if it can understand DSSM, even though it is character level trained, cant it correlate with glove ? Can you plz elaborate on this ?",
    "461746": "I gave the model in the paper a try, but it can't learn anything even at lower learning rates. If u apply the lr=0.05 as described in the paper, the loss overshoots. Using a lower lr makes the loss converge but the model doesn't learn anything of value.\n\nRead up on DSSM if you can, but from my understanding, the DSSM embedding conserves a temporal relational between the adjacent steps of the embedding layer outcome and hence the model is able to learn. \n\nYou can also try using Cov1D to get similar outcome, however in my experience it takes a long time to train it that way and wouldn't be feasible in this challenge.",
    "461860": "Oh! Ok. I will read about it"
  },
  "source": "meta"
}