{
  "id": 127350,
  "title": "Brief summary of 13th place solution (hide the pain Harold)",
  "url": "/competitions/tensorflow2-question-answering/writeups/trimorph-brief-summary-of-13th-place-solution-hide",
  "author_name": "",
  "post_date": "2020-02-18T13:27:34.507Z",
  "votes": 28,
  "comment_count": 6,
  "views": 0,
  "content": "<p>At first, congrats to every team that scored higher than we did - since we finished on 13th place, that means all of you finished in the gold zone.</p>\n\n<p>Our best solution consists of 4 bert-based models: 1 distilbert, 2 albert-large, and 1 bert-large WWM. </p>\n\n<p><a href=\"/kashnitsky\">@kashnitsky</a> has trained our best single model that scored around 0.64 on the local dev set (we used official NQ dev set for validation). He used original bert-joint repo and added hacks described in the <a href=\"https://arxiv.org/abs/1909.05286\">paper</a>. He also spent a lot of time trying to make ALBERT models work, but none of them (even xxlarge version) wasn't better than bert-large. </p>\n\n<p><a href=\"/yaroshevskiy\">@yaroshevskiy</a> implemented his own version of bert-joint in pytorch, including all the pre and post processing stuff. His best model is based on ALBERT large pretrained on squad 2.0. Oleg also came up with a trick that one might call \"window smoothing\". The trick addressed bad predictions of start/end probabilities on the window edges. The idea is that for those start/end logits that are close to the edge of the window we use a linear combination of the current window logits and logits from the neighboring window. This improved the score by around 0.01-0.02.</p>\n\n<p>I implemented my own pytorch model that is different from bert-joint in two aspects:\n- Instead of working on arbitrary chunked texts, I work on top of long answer candidates\n- start/end logits are predicted jointly by an attention-like layer, and the unrealistic start/end positions (like padding or question tokens) are filled with -inf</p>\n\n<p>Implementation for the start/end module is the following:</p>\n\n<p>```\nclass StartEndModule(nn.Module):</p>\n\n<pre><code>def __init__(self, input_dim, hidden_dim):\n    super().__init__()\n\n    self.start = nn.Conv1d(input_dim, hidden_dim, kernel_size=1)\n    self.end = nn.Conv1d(input_dim, hidden_dim, kernel_size=1)\n\ndef forward(self, hidden, text_mask):\n\n    start = self.start(hidden).unsqueeze(3)\n    end = self.end(hidden).unsqueeze(2)\n\n    logits = (start * end).sum(dim=1)\n    triu_mask = torch.triu(logits, diagonal=1) == 0\n    text_mask = ((text_mask.unsqueeze(2) * text_mask.unsqueeze(1)) &amp;lt; 0.5)\n    mask = text_mask | triu_mask\n    mask[:, 0, 0] = False\n    logits.masked_fill_(mask, float(\"-inf\"))\n\n    return logits, mask.float()\n</code></pre>\n\n<p>```</p>\n\n<p>Thus, logits is a square matrix where each entry is a score for a particular start/end pair. In order to find the best-scoring span, one just needs to compute an argmax over those scores.</p>\n\n<p>I also applied <a href=\"https://arxiv.org/abs/1803.05407\">SWA</a> to both of my models and got a nice boost in score (around 0.015), while Oleg and Yury reported none to minor improvements from SWA).</p>\n\n<p>In order to speedup inference, we used distilbert model for candidate prescoring. The idea is for all other models except for distilbert we ignore those windows/candidates that received low scores from the distilbert model.</p>\n\n<p>In order to blend our models together, we used a lightgbm boosting tree. For each candidate, we collect the corresponding scores from all the models as well as some meta-features (such as answer length or relative position of this candidate in the document) and the target is to predict if this candidate contains an answer. </p>\n\n<p>Our best blend achieved around 0.67 on the local dev set and 0.68 on the LB.</p>",
  "messages": [
    {
      "id": "727086",
      "postDate": "01/23/2020 12:57:12",
      "content": "<p>At first, congrats to every team that scored higher than we did - since we finished on 13th place, that means all of you finished in the gold zone.</p>\n\n<p>Our best solution consists of 4 bert-based models: 1 distilbert, 2 albert-large, and 1 bert-large WWM. </p>\n\n<p><a href=\"/kashnitsky\">@kashnitsky</a> has trained our best single model that scored around 0.64 on the local dev set (we used official NQ dev set for validation). He used original bert-joint repo and added hacks described in the <a href=\"https://arxiv.org/abs/1909.05286\">paper</a>. He also spent a lot of time trying to make ALBERT models work, but none of them (even xxlarge version) wasn't better than bert-large. </p>\n\n<p><a href=\"/yaroshevskiy\">@yaroshevskiy</a> implemented his own version of bert-joint in pytorch, including all the pre and post processing stuff. His best model is based on ALBERT large pretrained on squad 2.0. Oleg also came up with a trick that one might call \"window smoothing\". The trick addressed bad predictions of start/end probabilities on the window edges. The idea is that for those start/end logits that are close to the edge of the window we use a linear combination of the current window logits and logits from the neighboring window. This improved the score by around 0.01-0.02.</p>\n\n<p>I implemented my own pytorch model that is different from bert-joint in two aspects:\n- Instead of working on arbitrary chunked texts, I work on top of long answer candidates\n- start/end logits are predicted jointly by an attention-like layer, and the unrealistic start/end positions (like padding or question tokens) are filled with -inf</p>\n\n<p>Implementation for the start/end module is the following:</p>\n\n<p>```\nclass StartEndModule(nn.Module):</p>\n\n<pre><code>def __init__(self, input_dim, hidden_dim):\n    super().__init__()\n\n    self.start = nn.Conv1d(input_dim, hidden_dim, kernel_size=1)\n    self.end = nn.Conv1d(input_dim, hidden_dim, kernel_size=1)\n\ndef forward(self, hidden, text_mask):\n\n    start = self.start(hidden).unsqueeze(3)\n    end = self.end(hidden).unsqueeze(2)\n\n    logits = (start * end).sum(dim=1)\n    triu_mask = torch.triu(logits, diagonal=1) == 0\n    text_mask = ((text_mask.unsqueeze(2) * text_mask.unsqueeze(1)) &amp;lt; 0.5)\n    mask = text_mask | triu_mask\n    mask[:, 0, 0] = False\n    logits.masked_fill_(mask, float(\"-inf\"))\n\n    return logits, mask.float()\n</code></pre>\n\n<p>```</p>\n\n<p>Thus, logits is a square matrix where each entry is a score for a particular start/end pair. In order to find the best-scoring span, one just needs to compute an argmax over those scores.</p>\n\n<p>I also applied <a href=\"https://arxiv.org/abs/1803.05407\">SWA</a> to both of my models and got a nice boost in score (around 0.015), while Oleg and Yury reported none to minor improvements from SWA).</p>\n\n<p>In order to speedup inference, we used distilbert model for candidate prescoring. The idea is for all other models except for distilbert we ignore those windows/candidates that received low scores from the distilbert model.</p>\n\n<p>In order to blend our models together, we used a lightgbm boosting tree. For each candidate, we collect the corresponding scores from all the models as well as some meta-features (such as answer length or relative position of this candidate in the document) and the target is to predict if this candidate contains an answer. </p>\n\n<p>Our best blend achieved around 0.67 on the local dev set and 0.68 on the LB.</p>",
      "rawMarkdown": "At first, congrats to every team that scored higher than we did - since we finished on 13th place, that means all of you finished in the gold zone.\n\nOur best solution consists of 4 bert-based models: 1 distilbert, 2 albert-large, and 1 bert-large WWM. \n\n@kashnitsky has trained our best single model that scored around 0.64 on the local dev set (we used official NQ dev set for validation). He used original bert-joint repo and added hacks described in the [paper](https://arxiv.org/abs/1909.05286). He also spent a lot of time trying to make ALBERT models work, but none of them (even xxlarge version) wasn't better than bert-large. \n\n@yaroshevskiy implemented his own version of bert-joint in pytorch, including all the pre and post processing stuff. His best model is based on ALBERT large pretrained on squad 2.0. Oleg also came up with a trick that one might call \"window smoothing\". The trick addressed bad predictions of start/end probabilities on the window edges. The idea is that for those start/end logits that are close to the edge of the window we use a linear combination of the current window logits and logits from the neighboring window. This improved the score by around 0.01-0.02.\n\nI implemented my own pytorch model that is different from bert-joint in two aspects:\n- Instead of working on arbitrary chunked texts, I work on top of long answer candidates\n- start/end logits are predicted jointly by an attention-like layer, and the unrealistic start/end positions (like padding or question tokens) are filled with -inf\n\nImplementation for the start/end module is the following:\n\n```\nclass StartEndModule(nn.Module):\n\n    def __init__(self, input_dim, hidden_dim):\n        super().__init__()\n\n        self.start = nn.Conv1d(input_dim, hidden_dim, kernel_size=1)\n        self.end = nn.Conv1d(input_dim, hidden_dim, kernel_size=1)\n\n    def forward(self, hidden, text_mask):\n\n        start = self.start(hidden).unsqueeze(3)\n        end = self.end(hidden).unsqueeze(2)\n\n        logits = (start * end).sum(dim=1)\n        triu_mask = torch.triu(logits, diagonal=1) == 0\n        text_mask = ((text_mask.unsqueeze(2) * text_mask.unsqueeze(1)) &lt; 0.5)\n        mask = text_mask | triu_mask\n        mask[:, 0, 0] = False\n        logits.masked_fill_(mask, float(\"-inf\"))\n\n        return logits, mask.float()\n```\n\nThus, logits is a square matrix where each entry is a score for a particular start/end pair. In order to find the best-scoring span, one just needs to compute an argmax over those scores.\n\nI also applied [SWA](https://arxiv.org/abs/1803.05407) to both of my models and got a nice boost in score (around 0.015), while Oleg and Yury reported none to minor improvements from SWA).\n\nIn order to speedup inference, we used distilbert model for candidate prescoring. The idea is for all other models except for distilbert we ignore those windows/candidates that received low scores from the distilbert model.\n\nIn order to blend our models together, we used a lightgbm boosting tree. For each candidate, we collect the corresponding scores from all the models as well as some meta-features (such as answer length or relative position of this candidate in the document) and the target is to predict if this candidate contains an answer. \n\nOur best blend achieved around 0.67 on the local dev set and 0.68 on the LB.",
      "votes": null
    },
    {
      "id": "727179",
      "postDate": "01/23/2020 14:12:43",
      "content": "<p>Hi, thanks for your sharing! You implemented torch.optim.SGD with SWA method?</p>\n\n<p>Does SGD work for NLP problem?</p>",
      "rawMarkdown": "Hi, thanks for your sharing! You implemented torch.optim.SGD with SWA method?\n\nDoes SGD work for NLP problem?",
      "votes": null
    },
    {
      "id": "727181",
      "postDate": "01/23/2020 14:15:05",
      "content": "<p>I just averaged the weights. All the models were trained with Adam.</p>",
      "rawMarkdown": "I just averaged the weights. All the models were trained with Adam.",
      "votes": null
    },
    {
      "id": "727235",
      "postDate": "01/23/2020 14:59:39",
      "content": "<p>Sorry, I do not know much about SWA. If there is SWA, we can still implement scheduler?</p>",
      "rawMarkdown": "Sorry, I do not know much about SWA. If there is SWA, we can still implement scheduler?",
      "votes": null
    },
    {
      "id": "727438",
      "postDate": "01/23/2020 18:13:03",
      "content": "<p><a href=\"/xiaojiu1414\">@xiaojiu1414</a> yes you are free to use any LR scheduler. the idea of SWA is to average few trained model states - f.e. you save 10 checkpoints over last 10% of training and average them.</p>",
      "rawMarkdown": "xiaojiu1414 yes you are free to use any LR scheduler. the idea of SWA is to average few trained model states - f.e. you save 10 checkpoints over last 10% of training and average them.",
      "votes": null
    },
    {
      "id": "733315",
      "postDate": "01/31/2020 00:24:28",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "739430",
      "postDate": "02/07/2020 20:39:25",
      "content": "<p><a href=\"/ddanevskyi\">@ddanevskyi</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "rawMarkdown": "ddanevskyi, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 727179,
      "author_name": "xiaojiu1414",
      "author_url": "",
      "post_date": "01/23/2020 14:12:43",
      "content": "<p>Hi, thanks for your sharing! You implemented torch.optim.SGD with SWA method?</p>\n\n<p>Does SGD work for NLP problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 727181,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "01/23/2020 14:15:05",
          "content": "<p>I just averaged the weights. All the models were trained with Adam.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 727235,
          "author_name": "xiaojiu1414",
          "author_url": "",
          "post_date": "01/23/2020 14:59:39",
          "content": "<p>Sorry, I do not know much about SWA. If there is SWA, we can still implement scheduler?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 727438,
          "author_name": "yaroshevskiy",
          "author_url": "",
          "post_date": "01/23/2020 18:13:03",
          "content": "<p><a href=\"/xiaojiu1414\">@xiaojiu1414</a> yes you are free to use any LR scheduler. the idea of SWA is to average few trained model states - f.e. you save 10 checkpoints over last 10% of training and average them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 733315,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "01/31/2020 00:24:28",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 739430,
      "author_name": "vbmokin",
      "author_url": "",
      "post_date": "02/07/2020 20:39:25",
      "content": "<p><a href=\"/ddanevskyi\">@ddanevskyi</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "727086": "At first, congrats to every team that scored higher than we did - since we finished on 13th place, that means all of you finished in the gold zone.\n\nOur best solution consists of 4 bert-based models: 1 distilbert, 2 albert-large, and 1 bert-large WWM. \n\n@kashnitsky has trained our best single model that scored around 0.64 on the local dev set (we used official NQ dev set for validation). He used original bert-joint repo and added hacks described in the [paper](https://arxiv.org/abs/1909.05286). He also spent a lot of time trying to make ALBERT models work, but none of them (even xxlarge version) wasn't better than bert-large. \n\n@yaroshevskiy implemented his own version of bert-joint in pytorch, including all the pre and post processing stuff. His best model is based on ALBERT large pretrained on squad 2.0. Oleg also came up with a trick that one might call \"window smoothing\". The trick addressed bad predictions of start/end probabilities on the window edges. The idea is that for those start/end logits that are close to the edge of the window we use a linear combination of the current window logits and logits from the neighboring window. This improved the score by around 0.01-0.02.\n\nI implemented my own pytorch model that is different from bert-joint in two aspects:\n- Instead of working on arbitrary chunked texts, I work on top of long answer candidates\n- start/end logits are predicted jointly by an attention-like layer, and the unrealistic start/end positions (like padding or question tokens) are filled with -inf\n\nImplementation for the start/end module is the following:\n\n```\nclass StartEndModule(nn.Module):\n\n    def __init__(self, input_dim, hidden_dim):\n        super().__init__()\n\n        self.start = nn.Conv1d(input_dim, hidden_dim, kernel_size=1)\n        self.end = nn.Conv1d(input_dim, hidden_dim, kernel_size=1)\n\n    def forward(self, hidden, text_mask):\n\n        start = self.start(hidden).unsqueeze(3)\n        end = self.end(hidden).unsqueeze(2)\n\n        logits = (start * end).sum(dim=1)\n        triu_mask = torch.triu(logits, diagonal=1) == 0\n        text_mask = ((text_mask.unsqueeze(2) * text_mask.unsqueeze(1)) &lt; 0.5)\n        mask = text_mask | triu_mask\n        mask[:, 0, 0] = False\n        logits.masked_fill_(mask, float(\"-inf\"))\n\n        return logits, mask.float()\n```\n\nThus, logits is a square matrix where each entry is a score for a particular start/end pair. In order to find the best-scoring span, one just needs to compute an argmax over those scores.\n\nI also applied [SWA](https://arxiv.org/abs/1803.05407) to both of my models and got a nice boost in score (around 0.015), while Oleg and Yury reported none to minor improvements from SWA).\n\nIn order to speedup inference, we used distilbert model for candidate prescoring. The idea is for all other models except for distilbert we ignore those windows/candidates that received low scores from the distilbert model.\n\nIn order to blend our models together, we used a lightgbm boosting tree. For each candidate, we collect the corresponding scores from all the models as well as some meta-features (such as answer length or relative position of this candidate in the document) and the target is to predict if this candidate contains an answer. \n\nOur best blend achieved around 0.67 on the local dev set and 0.68 on the LB.",
    "727179": "Hi, thanks for your sharing! You implemented torch.optim.SGD with SWA method?\n\nDoes SGD work for NLP problem?",
    "727181": "I just averaged the weights. All the models were trained with Adam.",
    "727235": "Sorry, I do not know much about SWA. If there is SWA, we can still implement scheduler?",
    "727438": "xiaojiu1414 yes you are free to use any LR scheduler. the idea of SWA is to average few trained model states - f.e. you save 10 checkpoints over last 10% of training and average them.",
    "733315": "Thanks for sharing",
    "739430": "ddanevskyi, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)"
  },
  "source": "meta"
}