{
  "id": 127299,
  "title": "27th solution with luck and some questions from me",
  "url": "/competitions/tensorflow2-question-answering/discussion/127299",
  "author_name": "SchenbergZ",
  "post_date": "2020-01-23T08:08:09.599000",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks kaggle provides this competition and this is a big improvement for me to take an competition solo and reach this place.</p>\n\n<h2><strong>1. Wierd ROBERTA Achitecture</strong></h2>\n\n<p>Base from (<a href=\"https://github.com/bojone/bert4keras\">https://github.com/bojone/bert4keras</a>) My solution is based on a wierd ROBERTA structure inplemented in keras (<a href=\"https://www.kaggle.com/httpwwwfszyc/bert4keras4nq\">https://www.kaggle.com/httpwwwfszyc/bert4keras4nq</a>) as it is different from the huggingface in \n1. 2*1024 token_type_id_embedding layer (0 put pretrained weights, 1 set to np.zeros((1,1024)))\n2. no padding masking \n3. wierd zero masking in attention mask:\n```\n    def build(self):\n        x_in = Input(shape=(512, ), name='Input-Token')\n        s_in = Input(shape=(512, ), name='Input-Segment')\n        x, s = x_in, s_in</p>\n\n<pre><code>    sequence_mask = Lambda(lambda x: K.cast(K.greater(x, 0), 'float32'),\n                           name='Sequence-Mask')(x)\n\n    # Embedding\n    x = Embedding(input_dim=self.vocab_size,\n                  output_dim=self.embedding_size,\n                  embeddings_initializer=self.initializer,\n                  name='Embedding-Token')(x)\n    s = Embedding(input_dim=2, #1 or 2 , 2 finally because roberta need to train it\n                  output_dim=self.embedding_size,\n                  embeddings_initializer=self.initializer,\n                  name='Embedding-Segment')(s)\n    x = Add(name='Embedding-Token-Segment')([x, s])\n    if self.max_position_embeddings == 514:\n        x = RobertaPositionEmbeddings(input_dim=self.max_position_embeddings,\n                              output_dim=self.embedding_size,\n                              merge_mode='add',\n                              embeddings_initializer=self.initializer,\n                              name='Embedding-Position')([x,x_in])\n    else:\n        x = PositionEmbedding(input_dim=self.max_position_embeddings,\n                              output_dim=self.embedding_size,\n                              merge_mode='add',\n                              embeddings_initializer=self.initializer,\n                              name='Embedding-Position')(x)\n    x = LayerNormalization(name='Embedding-Norm')(x)\n    if self.dropout_rate &amp;gt; 0:\n        x = Dropout(rate=self.dropout_rate, name='Embedding-Dropout')(x)\n    if self.embedding_size != self.hidden_size:\n        x = Dense(units=self.hidden_size,\n                  kernel_initializer=self.initializer,\n                  name='Embedding-Mapping')(x)\n\n    layers = None\n    for i in range(self.num_hidden_layers):\n        attention_name = 'Encoder-%d-MultiHeadSelfAttention' % (i + 1)\n        feed_forward_name = 'Encoder-%d-FeedForward' % (i + 1)\n        x, layers = self.transformer_block(\n            inputs=x,\n            sequence_mask=sequence_mask,\n            attention_mask=self.compute_attention_mask(i, s_in),\n            attention_name=attention_name,\n            feed_forward_name=feed_forward_name,\n            input_layers=layers)\n        x = self.post_processing(i, x)\n        if not self.block_sharing:\n            layers = None\n\n    outputs = [x]\n</code></pre>\n\n<p>```\nI concatenate last 4 layers and put a single linear output for each output head. </p>\n\n<h2><strong>2. data distribution</strong></h2>\n\n<p>Samples-ratio of non-zero with 256 stride vs zero with stride 128 is 1:4.</p>\n\n<h2><strong>3. Training</strong></h2>\n\n<ol>\n<li>UseRadam with warmup 0.05 and train 1 epoch</li>\n<li>set different weights to match the distribution of dev set (I use 2 dev set for the 135000th to 140000th and for the 302373 to end section). As a result my loss is:</li>\n</ol>\n\n<p>*Total_loss = loss_weights1*sample_weights*start_loss+ loss_weights2*sample_weights*end_loss+ loss_weights3*sample_weights*answertype_loss*</p>\n\n<p>loss weights for [start, end, answer_type] is 1:1: (1/sampleweight.mean())</p>\n\n<h2><strong>4. Threshold killing False Positive</strong></h2>\n\n<p>The result of my solution is: CV: 0.478 because there are too many False Negative samples after I fix my metric error 10 days ago. So searching by my 2 devsets I finally choose a safe threshold ( the one slightly smaller than the threshold which reach max CV in order to lower the risk) . If a short answer score less than 0.5 or long answer smaller than 0.1 the answer will be blank.</p>\n\n<p>As I only upload 1 model, I use stride=128 for inference.\nMy result: CV 0.523, public LB 0.63, private LB 0.65</p>\n\n<h2><strong>5 My question</strong></h2>\n\n<ol>\n<li>how to mask padding:\nI tried to add a padding mask before embedding layer but which will raise error because layers after does not support mask... So I have to use this wierd ROBERTA architecture.</li>\n<li>why my attention mask never works:\nBase on time limit for me, when I realise there are some mistake on attention mask because of the code\n<code>\n    sequence_mask = Lambda(lambda x: K.cast(K.greater(x, 0), 'float32'),\n                           name='Sequence-Mask')(x)\n</code>\nI have no time to test it. So I just replace it by:\n<code>\n    sequence_mask = Lambda(lambda x: K.cast(K.not_equal(x, 1), 'float32'),\n                            name='Sequence-Mask')(x) \n</code>\nwhich replace token 1 because 1 represent padding. After TPU training I got 0.53 CV score. But when I plug this to gpu, the public LB only reach 0.48, which is wierd. So my question is: is my replacement of attention_mask really matching what I expect (mask token 1 of value matrix in attention layer)?</li>\n</ol>",
  "messages": [
    {
      "id": 726753,
      "postDate": "2020-01-23T08:08:09.600Z",
      "content": "<p>Thanks kaggle provides this competition and this is a big improvement for me to take an competition solo and reach this place.</p>\n\n<h2><strong>1. Wierd ROBERTA Achitecture</strong></h2>\n\n<p>Base from (<a href=\"https://github.com/bojone/bert4keras\">https://github.com/bojone/bert4keras</a>) My solution is based on a wierd ROBERTA structure inplemented in keras (<a href=\"https://www.kaggle.com/httpwwwfszyc/bert4keras4nq\">https://www.kaggle.com/httpwwwfszyc/bert4keras4nq</a>) as it is different from the huggingface in \n1. 2*1024 token_type_id_embedding layer (0 put pretrained weights, 1 set to np.zeros((1,1024)))\n2. no padding masking \n3. wierd zero masking in attention mask:\n```\n    def build(self):\n        x_in = Input(shape=(512, ), name='Input-Token')\n        s_in = Input(shape=(512, ), name='Input-Segment')\n        x, s = x_in, s_in</p>\n\n<pre><code>    sequence_mask = Lambda(lambda x: K.cast(K.greater(x, 0), 'float32'),\n                           name='Sequence-Mask')(x)\n\n    # Embedding\n    x = Embedding(input_dim=self.vocab_size,\n                  output_dim=self.embedding_size,\n                  embeddings_initializer=self.initializer,\n                  name='Embedding-Token')(x)\n    s = Embedding(input_dim=2, #1 or 2 , 2 finally because roberta need to train it\n                  output_dim=self.embedding_size,\n                  embeddings_initializer=self.initializer,\n                  name='Embedding-Segment')(s)\n    x = Add(name='Embedding-Token-Segment')([x, s])\n    if self.max_position_embeddings == 514:\n        x = RobertaPositionEmbeddings(input_dim=self.max_position_embeddings,\n                              output_dim=self.embedding_size,\n                              merge_mode='add',\n                              embeddings_initializer=self.initializer,\n                              name='Embedding-Position')([x,x_in])\n    else:\n        x = PositionEmbedding(input_dim=self.max_position_embeddings,\n                              output_dim=self.embedding_size,\n                              merge_mode='add',\n                              embeddings_initializer=self.initializer,\n                              name='Embedding-Position')(x)\n    x = LayerNormalization(name='Embedding-Norm')(x)\n    if self.dropout_rate &amp;gt; 0:\n        x = Dropout(rate=self.dropout_rate, name='Embedding-Dropout')(x)\n    if self.embedding_size != self.hidden_size:\n        x = Dense(units=self.hidden_size,\n                  kernel_initializer=self.initializer,\n                  name='Embedding-Mapping')(x)\n\n    layers = None\n    for i in range(self.num_hidden_layers):\n        attention_name = 'Encoder-%d-MultiHeadSelfAttention' % (i + 1)\n        feed_forward_name = 'Encoder-%d-FeedForward' % (i + 1)\n        x, layers = self.transformer_block(\n            inputs=x,\n            sequence_mask=sequence_mask,\n            attention_mask=self.compute_attention_mask(i, s_in),\n            attention_name=attention_name,\n            feed_forward_name=feed_forward_name,\n            input_layers=layers)\n        x = self.post_processing(i, x)\n        if not self.block_sharing:\n            layers = None\n\n    outputs = [x]\n</code></pre>\n\n<p>```\nI concatenate last 4 layers and put a single linear output for each output head. </p>\n\n<h2><strong>2. data distribution</strong></h2>\n\n<p>Samples-ratio of non-zero with 256 stride vs zero with stride 128 is 1:4.</p>\n\n<h2><strong>3. Training</strong></h2>\n\n<ol>\n<li>UseRadam with warmup 0.05 and train 1 epoch</li>\n<li>set different weights to match the distribution of dev set (I use 2 dev set for the 135000th to 140000th and for the 302373 to end section). As a result my loss is:</li>\n</ol>\n\n<p>*Total_loss = loss_weights1*sample_weights*start_loss+ loss_weights2*sample_weights*end_loss+ loss_weights3*sample_weights*answertype_loss*</p>\n\n<p>loss weights for [start, end, answer_type] is 1:1: (1/sampleweight.mean())</p>\n\n<h2><strong>4. Threshold killing False Positive</strong></h2>\n\n<p>The result of my solution is: CV: 0.478 because there are too many False Negative samples after I fix my metric error 10 days ago. So searching by my 2 devsets I finally choose a safe threshold ( the one slightly smaller than the threshold which reach max CV in order to lower the risk) . If a short answer score less than 0.5 or long answer smaller than 0.1 the answer will be blank.</p>\n\n<p>As I only upload 1 model, I use stride=128 for inference.\nMy result: CV 0.523, public LB 0.63, private LB 0.65</p>\n\n<h2><strong>5 My question</strong></h2>\n\n<ol>\n<li>how to mask padding:\nI tried to add a padding mask before embedding layer but which will raise error because layers after does not support mask... So I have to use this wierd ROBERTA architecture.</li>\n<li>why my attention mask never works:\nBase on time limit for me, when I realise there are some mistake on attention mask because of the code\n<code>\n    sequence_mask = Lambda(lambda x: K.cast(K.greater(x, 0), 'float32'),\n                           name='Sequence-Mask')(x)\n</code>\nI have no time to test it. So I just replace it by:\n<code>\n    sequence_mask = Lambda(lambda x: K.cast(K.not_equal(x, 1), 'float32'),\n                            name='Sequence-Mask')(x) \n</code>\nwhich replace token 1 because 1 represent padding. After TPU training I got 0.53 CV score. But when I plug this to gpu, the public LB only reach 0.48, which is wierd. So my question is: is my replacement of attention_mask really matching what I expect (mask token 1 of value matrix in attention layer)?</li>\n</ol>",
      "rawMarkdown": "Thanks kaggle provides this competition and this is a big improvement for me to take an competition solo and reach this place.\n\n## **1. Wierd ROBERTA Achitecture**\nBase from (https://github.com/bojone/bert4keras) My solution is based on a wierd ROBERTA structure inplemented in keras (https://www.kaggle.com/httpwwwfszyc/bert4keras4nq) as it is different from the huggingface in \n1. 2*1024 token_type_id_embedding layer (0 put pretrained weights, 1 set to np.zeros((1,1024)))\n2. no padding masking \n3. wierd zero masking in attention mask:\n```\n    def build(self):\n        x_in = Input(shape=(512, ), name='Input-Token')\n        s_in = Input(shape=(512, ), name='Input-Segment')\n        x, s = x_in, s_in\n\n        sequence_mask = Lambda(lambda x: K.cast(K.greater(x, 0), 'float32'),\n                               name='Sequence-Mask')(x)\n\n        # Embedding\n        x = Embedding(input_dim=self.vocab_size,\n                      output_dim=self.embedding_size,\n                      embeddings_initializer=self.initializer,\n                      name='Embedding-Token')(x)\n        s = Embedding(input_dim=2, #1 or 2 , 2 finally because roberta need to train it\n                      output_dim=self.embedding_size,\n                      embeddings_initializer=self.initializer,\n                      name='Embedding-Segment')(s)\n        x = Add(name='Embedding-Token-Segment')([x, s])\n        if self.max_position_embeddings == 514:\n            x = RobertaPositionEmbeddings(input_dim=self.max_position_embeddings,\n                                  output_dim=self.embedding_size,\n                                  merge_mode='add',\n                                  embeddings_initializer=self.initializer,\n                                  name='Embedding-Position')([x,x_in])\n        else:\n            x = PositionEmbedding(input_dim=self.max_position_embeddings,\n                                  output_dim=self.embedding_size,\n                                  merge_mode='add',\n                                  embeddings_initializer=self.initializer,\n                                  name='Embedding-Position')(x)\n        x = LayerNormalization(name='Embedding-Norm')(x)\n        if self.dropout_rate &gt; 0:\n            x = Dropout(rate=self.dropout_rate, name='Embedding-Dropout')(x)\n        if self.embedding_size != self.hidden_size:\n            x = Dense(units=self.hidden_size,\n                      kernel_initializer=self.initializer,\n                      name='Embedding-Mapping')(x)\n\n        layers = None\n        for i in range(self.num_hidden_layers):\n            attention_name = 'Encoder-%d-MultiHeadSelfAttention' % (i + 1)\n            feed_forward_name = 'Encoder-%d-FeedForward' % (i + 1)\n            x, layers = self.transformer_block(\n                inputs=x,\n                sequence_mask=sequence_mask,\n                attention_mask=self.compute_attention_mask(i, s_in),\n                attention_name=attention_name,\n                feed_forward_name=feed_forward_name,\n                input_layers=layers)\n            x = self.post_processing(i, x)\n            if not self.block_sharing:\n                layers = None\n\n        outputs = [x]\n```\nI concatenate last 4 layers and put a single linear output for each output head. \n## **2. data distribution**\nSamples-ratio of non-zero with 256 stride vs zero with stride 128 is 1:4.\n## **3. Training**\n1. UseRadam with warmup 0.05 and train 1 epoch\n2. set different weights to match the distribution of dev set (I use 2 dev set for the 135000th to 140000th and for the 302373 to end section). As a result my loss is:\n\n*Total\\_loss = loss_weights1\\*sample\\_weights\\*start\\_loss+ loss\\_weights2\\*sample\\_weights\\*end_loss+ loss\\_weights3\\*sample\\_weights\\*answertype\\_loss*\n\nloss weights for [start, end, answer_type] is 1:1: (1/sampleweight.mean())\n\n## **4. Threshold killing False Positive**\nThe result of my solution is: CV: 0.478 because there are too many False Negative samples after I fix my metric error 10 days ago. So searching by my 2 devsets I finally choose a safe threshold ( the one slightly smaller than the threshold which reach max CV in order to lower the risk) . If a short answer score less than 0.5 or long answer smaller than 0.1 the answer will be blank.\n\nAs I only upload 1 model, I use stride=128 for inference.\nMy result: CV 0.523, public LB 0.63, private LB 0.65\n\n## **5 My question**\n1. how to mask padding:\nI tried to add a padding mask before embedding layer but which will raise error because layers after does not support mask... So I have to use this wierd ROBERTA architecture.\n2. why my attention mask never works:\nBase on time limit for me, when I realise there are some mistake on attention mask because of the code\n```\n        sequence_mask = Lambda(lambda x: K.cast(K.greater(x, 0), 'float32'),\n                               name='Sequence-Mask')(x)\n```\nI have no time to test it. So I just replace it by:\n```\n        sequence_mask = Lambda(lambda x: K.cast(K.not_equal(x, 1), 'float32'),\n                                name='Sequence-Mask')(x) \n```\nwhich replace token 1 because 1 represent padding. After TPU training I got 0.53 CV score. But when I plug this to gpu, the public LB only reach 0.48, which is wierd. So my question is: is my replacement of attention_mask really matching what I expect (mask token 1 of value matrix in attention layer)?",
      "votes": 9
    },
    {
      "id": 726759,
      "postDate": "2020-01-23T08:13:25.777Z",
      "content": "<h2>Additions:</h2>\n\n<ol>\n<li>Because the private data seems close to the public, so my weighting strategy might be not good.</li>\n<li>RobertaPositionEmbedding comes from huggingface transformers</li>\n</ol>",
      "rawMarkdown": "## Additions: \n1. Because the private data seems close to the public, so my weighting strategy might be not good.\n2. RobertaPositionEmbedding comes from huggingface transformers"
    },
    {
      "id": 726779,
      "postDate": "2020-01-23T08:32:53.390Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 727834,
          "postDate": "2020-01-24T05:15:22.813Z",
          "content": "<p>You are welcome!</p>",
          "rawMarkdown": "You are welcome!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 726759,
      "author_name": "SchenbergZ",
      "author_url": "",
      "post_date": "2020-01-23T08:13:25.777000",
      "content": "<h2>Additions:</h2>\n\n<ol>\n<li>Because the private data seems close to the public, so my weighting strategy might be not good.</li>\n<li>RobertaPositionEmbedding comes from huggingface transformers</li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 726779,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-23T08:32:53.390000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 727834,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2020-01-24T05:15:22.813000",
          "content": "<p>You are welcome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "726753": "Thanks kaggle provides this competition and this is a big improvement for me to take an competition solo and reach this place.\n\n## **1. Wierd ROBERTA Achitecture**\nBase from (https://github.com/bojone/bert4keras) My solution is based on a wierd ROBERTA structure inplemented in keras (https://www.kaggle.com/httpwwwfszyc/bert4keras4nq) as it is different from the huggingface in \n1. 2*1024 token_type_id_embedding layer (0 put pretrained weights, 1 set to np.zeros((1,1024)))\n2. no padding masking \n3. wierd zero masking in attention mask:\n```\n    def build(self):\n        x_in = Input(shape=(512, ), name='Input-Token')\n        s_in = Input(shape=(512, ), name='Input-Segment')\n        x, s = x_in, s_in\n\n        sequence_mask = Lambda(lambda x: K.cast(K.greater(x, 0), 'float32'),\n                               name='Sequence-Mask')(x)\n\n        # Embedding\n        x = Embedding(input_dim=self.vocab_size,\n                      output_dim=self.embedding_size,\n                      embeddings_initializer=self.initializer,\n                      name='Embedding-Token')(x)\n        s = Embedding(input_dim=2, #1 or 2 , 2 finally because roberta need to train it\n                      output_dim=self.embedding_size,\n                      embeddings_initializer=self.initializer,\n                      name='Embedding-Segment')(s)\n        x = Add(name='Embedding-Token-Segment')([x, s])\n        if self.max_position_embeddings == 514:\n            x = RobertaPositionEmbeddings(input_dim=self.max_position_embeddings,\n                                  output_dim=self.embedding_size,\n                                  merge_mode='add',\n                                  embeddings_initializer=self.initializer,\n                                  name='Embedding-Position')([x,x_in])\n        else:\n            x = PositionEmbedding(input_dim=self.max_position_embeddings,\n                                  output_dim=self.embedding_size,\n                                  merge_mode='add',\n                                  embeddings_initializer=self.initializer,\n                                  name='Embedding-Position')(x)\n        x = LayerNormalization(name='Embedding-Norm')(x)\n        if self.dropout_rate &gt; 0:\n            x = Dropout(rate=self.dropout_rate, name='Embedding-Dropout')(x)\n        if self.embedding_size != self.hidden_size:\n            x = Dense(units=self.hidden_size,\n                      kernel_initializer=self.initializer,\n                      name='Embedding-Mapping')(x)\n\n        layers = None\n        for i in range(self.num_hidden_layers):\n            attention_name = 'Encoder-%d-MultiHeadSelfAttention' % (i + 1)\n            feed_forward_name = 'Encoder-%d-FeedForward' % (i + 1)\n            x, layers = self.transformer_block(\n                inputs=x,\n                sequence_mask=sequence_mask,\n                attention_mask=self.compute_attention_mask(i, s_in),\n                attention_name=attention_name,\n                feed_forward_name=feed_forward_name,\n                input_layers=layers)\n            x = self.post_processing(i, x)\n            if not self.block_sharing:\n                layers = None\n\n        outputs = [x]\n```\nI concatenate last 4 layers and put a single linear output for each output head. \n## **2. data distribution**\nSamples-ratio of non-zero with 256 stride vs zero with stride 128 is 1:4.\n## **3. Training**\n1. UseRadam with warmup 0.05 and train 1 epoch\n2. set different weights to match the distribution of dev set (I use 2 dev set for the 135000th to 140000th and for the 302373 to end section). As a result my loss is:\n\n*Total\\_loss = loss_weights1\\*sample\\_weights\\*start\\_loss+ loss\\_weights2\\*sample\\_weights\\*end_loss+ loss\\_weights3\\*sample\\_weights\\*answertype\\_loss*\n\nloss weights for [start, end, answer_type] is 1:1: (1/sampleweight.mean())\n\n## **4. Threshold killing False Positive**\nThe result of my solution is: CV: 0.478 because there are too many False Negative samples after I fix my metric error 10 days ago. So searching by my 2 devsets I finally choose a safe threshold ( the one slightly smaller than the threshold which reach max CV in order to lower the risk) . If a short answer score less than 0.5 or long answer smaller than 0.1 the answer will be blank.\n\nAs I only upload 1 model, I use stride=128 for inference.\nMy result: CV 0.523, public LB 0.63, private LB 0.65\n\n## **5 My question**\n1. how to mask padding:\nI tried to add a padding mask before embedding layer but which will raise error because layers after does not support mask... So I have to use this wierd ROBERTA architecture.\n2. why my attention mask never works:\nBase on time limit for me, when I realise there are some mistake on attention mask because of the code\n```\n        sequence_mask = Lambda(lambda x: K.cast(K.greater(x, 0), 'float32'),\n                               name='Sequence-Mask')(x)\n```\nI have no time to test it. So I just replace it by:\n```\n        sequence_mask = Lambda(lambda x: K.cast(K.not_equal(x, 1), 'float32'),\n                                name='Sequence-Mask')(x) \n```\nwhich replace token 1 because 1 represent padding. After TPU training I got 0.53 CV score. But when I plug this to gpu, the public LB only reach 0.48, which is wierd. So my question is: is my replacement of attention_mask really matching what I expect (mask token 1 of value matrix in attention layer)?",
    "726759": "## Additions: \n1. Because the private data seems close to the public, so my weighting strategy might be not good.\n2. RobertaPositionEmbedding comes from huggingface transformers",
    "726779": ""
  }
}