{
  "id": 127232,
  "title": "31st solution with custom loss ",
  "url": "/competitions/tensorflow2-question-answering/writeups/higepon-31st-solution-with-custom-loss",
  "author_name": "",
  "post_date": "2020-01-23T01:26:58.121643600Z",
  "votes": 22,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Thank you Kaggle and Kaggle community for this awesome competition. I learned a lot.\nI wasn’t able to do almost anything the last two weeks due to my personal reason, but it has been really fun.</p>\n\n<h2>My model</h2>\n\n<ul>\n<li>Public 0.68 Public 0.65</li>\n<li>Single PyTorch Bert model</li>\n<li>fine-tune bert-large-uncased-whole-word-masking-finetuned-squad for 1 epoch. \n<ul><li>2 epochs got better Private 0.68 Public 0.65 but I didn't choose it :(</li></ul></li>\n<li>learning rate 3e-5 instead of 5e-5</li>\n<li>Down sampled null instance training data.</li>\n<li>Penalize training data with answer in stride in loss function.</li>\n<li>Simply removed HTML tags</li>\n<li>Parameters search using short/long score.</li>\n</ul>\n\n<h2>down sampling</h2>\n\n<p><code>\nflattened_examples = list(itertools.chain.from_iterable(examples))\nnull_instances = []\nannotated_instances = []\nfor e in flattened_examples:\n    if e.class_label == 'unknown':\n        null_instances.append(e)\n    else:\n        annotated_instances.append(e)\nlen_null = len(null_instances)\nlen_downsampled = int(len_null / 50) if len_null &gt; 50 else 0\ndownsampled = random.sample(null_instances, len_downsampled)\nlogging.info('    down sampling nonnull(%d) null(%d) to null(%d)', len(annotated_instances), len_null, len(downsampled))\nself.examples = downsampled + annotated_instanceCan someone share sgse\n</code></p>\n\n<h2>loss function</h2>\n\n<p>```\ndef loss_fn(preds, labels, no_answers):</p>\n\n<pre><code>start_preds, end_preds, class_preds = preds\nstart_labels, end_labels, class_labels = labels\n\nhas_answers = [not x for x in no_answers]\n\nstart_preds_no_answer = start_preds[no_answers]\nstart_preds_has_answer = start_preds[has_answers]\nend_preds_no_answer = end_preds[no_answers]\nend_preds_has_answer = end_preds[has_answers]\nclass_preds_no_answer = class_preds[no_answers]\nclass_preds_has_answer = class_preds[has_answers]\nstart_labels_no_answer = start_labels[no_answers]\nstart_labels_has_answer = start_labels[has_answers]\nend_labels_no_answer = end_labels[no_answers]\nend_labels_has_answer = end_labels[has_answers]\nclass_labels_no_answer = class_labels[no_answers]\nclass_labels_has_answer = class_labels[has_answers]\n\nloss_no_answer = 0\nloss_has_answer = 0\n# has answer\nif len(start_preds_has_answer) &gt; 0:\n    start_loss = nn.CrossEntropyLoss(ignore_index=-1)(start_preds_has_answer, start_labels_has_answer)\n    end_loss = nn.CrossEntropyLoss(ignore_index=-1)(end_preds_has_answer, end_labels_has_answer)\n    class_loss = nn.CrossEntropyLoss()(class_preds_has_answer, class_labels_has_answer)\n    loss_has_answer = start_loss + end_loss + class_loss\n\nif len(start_preds_no_answer) &gt; 0:\n    start_loss = nn.CrossEntropyLoss(ignore_index=-1)(start_preds_no_answer, start_labels_no_answer)\n    end_loss = nn.CrossEntropyLoss(ignore_index=-1)(end_preds_no_answer, end_labels_no_answer)\n    class_loss = nn.CrossEntropyLoss()(class_preds_no_answer, class_labels_no_answer)\n    loss_no_answer = start_loss + end_loss + class_loss\n\nreturn loss_has_answer * 2 + loss_no_answer\n</code></pre>\n\n<p>```</p>\n\n<h2>What I didn't try</h2>\n\n<ul>\n<li>p/table tag annotations</li>\n<li>TPU</li>\n<li>more post processing</li>\n</ul>",
  "messages": [
    {
      "id": "726360",
      "postDate": "01/23/2020 01:26:58",
      "content": "<p>Thank you Kaggle and Kaggle community for this awesome competition. I learned a lot.\nI wasn’t able to do almost anything the last two weeks due to my personal reason, but it has been really fun.</p>\n\n<h2>My model</h2>\n\n<ul>\n<li>Public 0.68 Public 0.65</li>\n<li>Single PyTorch Bert model</li>\n<li>fine-tune bert-large-uncased-whole-word-masking-finetuned-squad for 1 epoch. \n<ul><li>2 epochs got better Private 0.68 Public 0.65 but I didn't choose it :(</li></ul></li>\n<li>learning rate 3e-5 instead of 5e-5</li>\n<li>Down sampled null instance training data.</li>\n<li>Penalize training data with answer in stride in loss function.</li>\n<li>Simply removed HTML tags</li>\n<li>Parameters search using short/long score.</li>\n</ul>\n\n<h2>down sampling</h2>\n\n<p><code>\nflattened_examples = list(itertools.chain.from_iterable(examples))\nnull_instances = []\nannotated_instances = []\nfor e in flattened_examples:\n    if e.class_label == 'unknown':\n        null_instances.append(e)\n    else:\n        annotated_instances.append(e)\nlen_null = len(null_instances)\nlen_downsampled = int(len_null / 50) if len_null &gt; 50 else 0\ndownsampled = random.sample(null_instances, len_downsampled)\nlogging.info('    down sampling nonnull(%d) null(%d) to null(%d)', len(annotated_instances), len_null, len(downsampled))\nself.examples = downsampled + annotated_instanceCan someone share sgse\n</code></p>\n\n<h2>loss function</h2>\n\n<p>```\ndef loss_fn(preds, labels, no_answers):</p>\n\n<pre><code>start_preds, end_preds, class_preds = preds\nstart_labels, end_labels, class_labels = labels\n\nhas_answers = [not x for x in no_answers]\n\nstart_preds_no_answer = start_preds[no_answers]\nstart_preds_has_answer = start_preds[has_answers]\nend_preds_no_answer = end_preds[no_answers]\nend_preds_has_answer = end_preds[has_answers]\nclass_preds_no_answer = class_preds[no_answers]\nclass_preds_has_answer = class_preds[has_answers]\nstart_labels_no_answer = start_labels[no_answers]\nstart_labels_has_answer = start_labels[has_answers]\nend_labels_no_answer = end_labels[no_answers]\nend_labels_has_answer = end_labels[has_answers]\nclass_labels_no_answer = class_labels[no_answers]\nclass_labels_has_answer = class_labels[has_answers]\n\nloss_no_answer = 0\nloss_has_answer = 0\n# has answer\nif len(start_preds_has_answer) &gt; 0:\n    start_loss = nn.CrossEntropyLoss(ignore_index=-1)(start_preds_has_answer, start_labels_has_answer)\n    end_loss = nn.CrossEntropyLoss(ignore_index=-1)(end_preds_has_answer, end_labels_has_answer)\n    class_loss = nn.CrossEntropyLoss()(class_preds_has_answer, class_labels_has_answer)\n    loss_has_answer = start_loss + end_loss + class_loss\n\nif len(start_preds_no_answer) &gt; 0:\n    start_loss = nn.CrossEntropyLoss(ignore_index=-1)(start_preds_no_answer, start_labels_no_answer)\n    end_loss = nn.CrossEntropyLoss(ignore_index=-1)(end_preds_no_answer, end_labels_no_answer)\n    class_loss = nn.CrossEntropyLoss()(class_preds_no_answer, class_labels_no_answer)\n    loss_no_answer = start_loss + end_loss + class_loss\n\nreturn loss_has_answer * 2 + loss_no_answer\n</code></pre>\n\n<p>```</p>\n\n<h2>What I didn't try</h2>\n\n<ul>\n<li>p/table tag annotations</li>\n<li>TPU</li>\n<li>more post processing</li>\n</ul>",
      "rawMarkdown": "Thank you Kaggle and Kaggle community for this awesome competition. I learned a lot.\nI wasn’t able to do almost anything the last two weeks due to my personal reason, but it has been really fun.\n\n## My model\n- Public 0.68 Public 0.65\n- Single PyTorch Bert model\n- fine-tune bert-large-uncased-whole-word-masking-finetuned-squad for 1 epoch. \n    - 2 epochs got better Private 0.68 Public 0.65 but I didn't choose it :(\n- learning rate 3e-5 instead of 5e-5\n- Down sampled null instance training data.\n- Penalize training data with answer in stride in loss function.\n- Simply removed HTML tags\n- Parameters search using short/long score.\n\n## down sampling\n```\nflattened_examples = list(itertools.chain.from_iterable(examples))\nnull_instances = []\nannotated_instances = []\nfor e in flattened_examples:\n    if e.class_label == 'unknown':\n        null_instances.append(e)\n    else:\n        annotated_instances.append(e)\nlen_null = len(null_instances)\nlen_downsampled = int(len_null / 50) if len_null &gt; 50 else 0\ndownsampled = random.sample(null_instances, len_downsampled)\nlogging.info('    down sampling nonnull(%d) null(%d) to null(%d)', len(annotated_instances), len_null, len(downsampled))\nself.examples = downsampled + annotated_instanceCan someone share sgse\n```\n \n## loss function\n```\ndef loss_fn(preds, labels, no_answers):\n\n    start_preds, end_preds, class_preds = preds\n    start_labels, end_labels, class_labels = labels\n    \n    has_answers = [not x for x in no_answers]\n    \n    start_preds_no_answer = start_preds[no_answers]\n    start_preds_has_answer = start_preds[has_answers]\n    end_preds_no_answer = end_preds[no_answers]\n    end_preds_has_answer = end_preds[has_answers]\n    class_preds_no_answer = class_preds[no_answers]\n    class_preds_has_answer = class_preds[has_answers]\n    start_labels_no_answer = start_labels[no_answers]\n    start_labels_has_answer = start_labels[has_answers]\n    end_labels_no_answer = end_labels[no_answers]\n    end_labels_has_answer = end_labels[has_answers]\n    class_labels_no_answer = class_labels[no_answers]\n    class_labels_has_answer = class_labels[has_answers]\n\n    loss_no_answer = 0\n    loss_has_answer = 0\n    # has answer\n    if len(start_preds_has_answer) &gt; 0:\n        start_loss = nn.CrossEntropyLoss(ignore_index=-1)(start_preds_has_answer, start_labels_has_answer)\n        end_loss = nn.CrossEntropyLoss(ignore_index=-1)(end_preds_has_answer, end_labels_has_answer)\n        class_loss = nn.CrossEntropyLoss()(class_preds_has_answer, class_labels_has_answer)\n        loss_has_answer = start_loss + end_loss + class_loss\n\n    if len(start_preds_no_answer) &gt; 0:\n        start_loss = nn.CrossEntropyLoss(ignore_index=-1)(start_preds_no_answer, start_labels_no_answer)\n        end_loss = nn.CrossEntropyLoss(ignore_index=-1)(end_preds_no_answer, end_labels_no_answer)\n        class_loss = nn.CrossEntropyLoss()(class_preds_no_answer, class_labels_no_answer)\n        loss_no_answer = start_loss + end_loss + class_loss\n        \n    return loss_has_answer * 2 + loss_no_answer\n```\n\n## What I didn't try\n- p/table tag annotations\n- TPU\n- more post processing",
      "votes": null
    },
    {
      "id": "726400",
      "postDate": "01/23/2020 01:56:23",
      "content": "<p>Very interesting! If possible, please share your model training code as well. It would be greatly appreciated!</p>",
      "rawMarkdown": "Very interesting! If possible, please share your model training code as well. It would be greatly appreciated!",
      "votes": null
    },
    {
      "id": "726503",
      "postDate": "01/23/2020 03:18:52",
      "content": "<p>Congratulations\nThanks for sharing your Approach &amp; Insights!! <a href=\"/higepon\">@higepon</a> </p>",
      "rawMarkdown": "Congratulations\nThanks for sharing your Approach &amp; Insights!! @higepon",
      "votes": null
    },
    {
      "id": "726579",
      "postDate": "01/23/2020 05:13:56",
      "content": "<p>Congratulations <a href=\"/higepon\">@higepon</a>! Could you share the motivation/intuition to penalize training data with answer in stride in loss function?</p>",
      "rawMarkdown": "Congratulations @higepon! Could you share the motivation/intuition to penalize training data with answer in stride in loss function?",
      "votes": null
    },
    {
      "id": "726587",
      "postDate": "01/23/2020 05:30:42",
      "content": "<p>I borrowed the idea from <a href=\"https://web.stanford.edu/class/cs224n/reports/default/15812785.pdf\">https://web.stanford.edu/class/cs224n/reports/default/15812785.pdf</a> section 3.2.2😀</p>",
      "rawMarkdown": "I borrowed the idea from https://web.stanford.edu/class/cs224n/reports/default/15812785.pdf section 3.2.2😀",
      "votes": null
    },
    {
      "id": "726605",
      "postDate": "01/23/2020 05:51:15",
      "content": "<p>Congratulations! <a href=\"/higepon\">@higepon</a> You have been super active thoughout the competition.\nI wanted to ask, did you try to play with long/short score? Like score being a simple sum of start and end logit or harmonic mean?</p>",
      "rawMarkdown": "Congratulations! @higepon You have been super active thoughout the competition.\nI wanted to ask, did you try to play with long/short score? Like score being a simple sum of start and end logit or harmonic mean?",
      "votes": null
    },
    {
      "id": "726612",
      "postDate": "01/23/2020 06:00:20",
      "content": "<p>Thank you! I didn’t. I just used the original definition of the scores. Maybe I should have done that. Hope some of top teams can share their thoughts on it :)</p>",
      "rawMarkdown": "Thank you! I didn’t. I just used the original definition of the scores. Maybe I should have done that. Hope some of top teams can share their thoughts on it :)",
      "votes": null
    },
    {
      "id": "726649",
      "postDate": "01/23/2020 06:28:54",
      "content": "<p>Congratulations <a href=\"/higepon\">@higepon</a> . Would you be open to sharing your pytorch code?</p>",
      "rawMarkdown": "Congratulations @higepon . Would you be open to sharing your pytorch code?",
      "votes": null
    },
    {
      "id": "726928",
      "postDate": "01/23/2020 10:18:10",
      "content": "<p>Congrats &amp; Thanks for sharing!!🎉 😄 👍 </p>",
      "rawMarkdown": "Congrats &amp; Thanks for sharing!!🎉 😄 👍",
      "votes": null
    },
    {
      "id": "726937",
      "postDate": "01/23/2020 10:21:44",
      "content": "<p>Congrats <a href=\"/higepon\">@higepon</a> !!\nYour posts in forum help my understanding and your score really enhanced my motivation, Thanks a lot!\n1 question) What you jumped in LB from 0.60 to 0.65 is the idea of custom loss? </p>",
      "rawMarkdown": "Congrats @higepon !!\nYour posts in forum help my understanding and your score really enhanced my motivation, Thanks a lot!\n1 question) What you jumped in LB from 0.60 to 0.65 is the idea of custom loss?",
      "votes": null
    },
    {
      "id": "727055",
      "postDate": "01/23/2020 12:23:07",
      "content": "<p>Likewise.\nThe jump was a combination of the custom and down sampling.</p>",
      "rawMarkdown": "Likewise.\nThe jump was a combination of the custom and down sampling.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 726400,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "01/23/2020 01:56:23",
      "content": "<p>Very interesting! If possible, please share your model training code as well. It would be greatly appreciated!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 726503,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "01/23/2020 03:18:52",
      "content": "<p>Congratulations\nThanks for sharing your Approach &amp; Insights!! <a href=\"/higepon\">@higepon</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 726579,
      "author_name": "thedrcat",
      "author_url": "",
      "post_date": "01/23/2020 05:13:56",
      "content": "<p>Congratulations <a href=\"/higepon\">@higepon</a>! Could you share the motivation/intuition to penalize training data with answer in stride in loss function?</p>",
      "votes": null,
      "replies": [
        {
          "id": 726587,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/23/2020 05:30:42",
          "content": "<p>I borrowed the idea from <a href=\"https://web.stanford.edu/class/cs224n/reports/default/15812785.pdf\">https://web.stanford.edu/class/cs224n/reports/default/15812785.pdf</a> section 3.2.2😀</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 726605,
      "author_name": "rajraviprajapat",
      "author_url": "",
      "post_date": "01/23/2020 05:51:15",
      "content": "<p>Congratulations! <a href=\"/higepon\">@higepon</a> You have been super active thoughout the competition.\nI wanted to ask, did you try to play with long/short score? Like score being a simple sum of start and end logit or harmonic mean?</p>",
      "votes": null,
      "replies": [
        {
          "id": 726612,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/23/2020 06:00:20",
          "content": "<p>Thank you! I didn’t. I just used the original definition of the scores. Maybe I should have done that. Hope some of top teams can share their thoughts on it :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 726649,
      "author_name": "axel81",
      "author_url": "",
      "post_date": "01/23/2020 06:28:54",
      "content": "<p>Congratulations <a href=\"/higepon\">@higepon</a> . Would you be open to sharing your pytorch code?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 726928,
      "author_name": "mashlyn",
      "author_url": "",
      "post_date": "01/23/2020 10:18:10",
      "content": "<p>Congrats &amp; Thanks for sharing!!🎉 😄 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 726937,
      "author_name": "kentaronakanishi",
      "author_url": "",
      "post_date": "01/23/2020 10:21:44",
      "content": "<p>Congrats <a href=\"/higepon\">@higepon</a> !!\nYour posts in forum help my understanding and your score really enhanced my motivation, Thanks a lot!\n1 question) What you jumped in LB from 0.60 to 0.65 is the idea of custom loss? </p>",
      "votes": null,
      "replies": [
        {
          "id": 727055,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "01/23/2020 12:23:07",
          "content": "<p>Likewise.\nThe jump was a combination of the custom and down sampling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "726360": "Thank you Kaggle and Kaggle community for this awesome competition. I learned a lot.\nI wasn’t able to do almost anything the last two weeks due to my personal reason, but it has been really fun.\n\n## My model\n- Public 0.68 Public 0.65\n- Single PyTorch Bert model\n- fine-tune bert-large-uncased-whole-word-masking-finetuned-squad for 1 epoch. \n    - 2 epochs got better Private 0.68 Public 0.65 but I didn't choose it :(\n- learning rate 3e-5 instead of 5e-5\n- Down sampled null instance training data.\n- Penalize training data with answer in stride in loss function.\n- Simply removed HTML tags\n- Parameters search using short/long score.\n\n## down sampling\n```\nflattened_examples = list(itertools.chain.from_iterable(examples))\nnull_instances = []\nannotated_instances = []\nfor e in flattened_examples:\n    if e.class_label == 'unknown':\n        null_instances.append(e)\n    else:\n        annotated_instances.append(e)\nlen_null = len(null_instances)\nlen_downsampled = int(len_null / 50) if len_null &gt; 50 else 0\ndownsampled = random.sample(null_instances, len_downsampled)\nlogging.info('    down sampling nonnull(%d) null(%d) to null(%d)', len(annotated_instances), len_null, len(downsampled))\nself.examples = downsampled + annotated_instanceCan someone share sgse\n```\n \n## loss function\n```\ndef loss_fn(preds, labels, no_answers):\n\n    start_preds, end_preds, class_preds = preds\n    start_labels, end_labels, class_labels = labels\n    \n    has_answers = [not x for x in no_answers]\n    \n    start_preds_no_answer = start_preds[no_answers]\n    start_preds_has_answer = start_preds[has_answers]\n    end_preds_no_answer = end_preds[no_answers]\n    end_preds_has_answer = end_preds[has_answers]\n    class_preds_no_answer = class_preds[no_answers]\n    class_preds_has_answer = class_preds[has_answers]\n    start_labels_no_answer = start_labels[no_answers]\n    start_labels_has_answer = start_labels[has_answers]\n    end_labels_no_answer = end_labels[no_answers]\n    end_labels_has_answer = end_labels[has_answers]\n    class_labels_no_answer = class_labels[no_answers]\n    class_labels_has_answer = class_labels[has_answers]\n\n    loss_no_answer = 0\n    loss_has_answer = 0\n    # has answer\n    if len(start_preds_has_answer) &gt; 0:\n        start_loss = nn.CrossEntropyLoss(ignore_index=-1)(start_preds_has_answer, start_labels_has_answer)\n        end_loss = nn.CrossEntropyLoss(ignore_index=-1)(end_preds_has_answer, end_labels_has_answer)\n        class_loss = nn.CrossEntropyLoss()(class_preds_has_answer, class_labels_has_answer)\n        loss_has_answer = start_loss + end_loss + class_loss\n\n    if len(start_preds_no_answer) &gt; 0:\n        start_loss = nn.CrossEntropyLoss(ignore_index=-1)(start_preds_no_answer, start_labels_no_answer)\n        end_loss = nn.CrossEntropyLoss(ignore_index=-1)(end_preds_no_answer, end_labels_no_answer)\n        class_loss = nn.CrossEntropyLoss()(class_preds_no_answer, class_labels_no_answer)\n        loss_no_answer = start_loss + end_loss + class_loss\n        \n    return loss_has_answer * 2 + loss_no_answer\n```\n\n## What I didn't try\n- p/table tag annotations\n- TPU\n- more post processing",
    "726400": "Very interesting! If possible, please share your model training code as well. It would be greatly appreciated!",
    "726503": "Congratulations\nThanks for sharing your Approach &amp; Insights!! @higepon",
    "726579": "Congratulations @higepon! Could you share the motivation/intuition to penalize training data with answer in stride in loss function?",
    "726587": "I borrowed the idea from https://web.stanford.edu/class/cs224n/reports/default/15812785.pdf section 3.2.2😀",
    "726605": "Congratulations! @higepon You have been super active thoughout the competition.\nI wanted to ask, did you try to play with long/short score? Like score being a simple sum of start and end logit or harmonic mean?",
    "726612": "Thank you! I didn’t. I just used the original definition of the scores. Maybe I should have done that. Hope some of top teams can share their thoughts on it :)",
    "726649": "Congratulations @higepon . Would you be open to sharing your pytorch code?",
    "726928": "Congrats &amp; Thanks for sharing!!🎉 😄 👍",
    "726937": "Congrats @higepon !!\nYour posts in forum help my understanding and your score really enhanced my motivation, Thanks a lot!\n1 question) What you jumped in LB from 0.60 to 0.65 is the idea of custom loss?",
    "727055": "Likewise.\nThe jump was a combination of the custom and down sampling."
  },
  "source": "meta"
}