{
  "id": 119992,
  "title": "Loss used in BertForQuestionAnswering",
  "url": "/competitions/tensorflow2-question-answering/discussion/119992",
  "author_name": "",
  "post_date": "2019-12-03T00:20:55.051261600Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have checked the model used in huggingface github for QuestionAnswering tasks and I  have found this (<a href=\"https://github.com/huggingface/transformers/blob/master/transformers/modeling_bert.py\">link</a>)\n```\nloss_fct = CrossEntropyLoss(ignore_index=ignored_index)</p>\n\n<p>start_loss = loss_fct(start_logits, start_positions)</p>\n\n<p>end_loss = loss_fct(end_logits, end_positions)</p>\n\n<p>total_loss = (start_loss + end_loss) / 2</p>\n\n<p>```\nIt is my first NLP competition and I  can't understand why the loss used here is CrossEntropyLoss. \nCan someone explain why we're using the CE loss instead of MSE? </p>",
  "messages": [
    {
      "id": "686246",
      "postDate": "12/03/2019 00:20:55",
      "content": "<p>I have checked the model used in huggingface github for QuestionAnswering tasks and I  have found this (<a href=\"https://github.com/huggingface/transformers/blob/master/transformers/modeling_bert.py\">link</a>)\n```\nloss_fct = CrossEntropyLoss(ignore_index=ignored_index)</p>\n\n<p>start_loss = loss_fct(start_logits, start_positions)</p>\n\n<p>end_loss = loss_fct(end_logits, end_positions)</p>\n\n<p>total_loss = (start_loss + end_loss) / 2</p>\n\n<p>```\nIt is my first NLP competition and I  can't understand why the loss used here is CrossEntropyLoss. \nCan someone explain why we're using the CE loss instead of MSE? </p>",
      "rawMarkdown": "I have checked the model used in huggingface github for QuestionAnswering tasks and I  have found this ([link](https://github.com/huggingface/transformers/blob/master/transformers/modeling_bert.py))\n```\nloss_fct = CrossEntropyLoss(ignore_index=ignored_index)\n\nstart_loss = loss_fct(start_logits, start_positions)\n\nend_loss = loss_fct(end_logits, end_positions)\n\ntotal_loss = (start_loss + end_loss) / 2\n\n```\nIt is my first NLP competition and I  can't understand why the loss used here is CrossEntropyLoss. \nCan someone explain why we're using the CE loss instead of MSE?",
      "votes": null
    },
    {
      "id": "686304",
      "postDate": "12/03/2019 02:17:57",
      "content": "<p>The most used label start_positions  maybe 5,10,100, if you change this to one-hot format[0,0,0,....1...,0] ,you can also use MSE loss</p>",
      "rawMarkdown": "The most used label start_positions  maybe 5,10,100, if you change this to one-hot format[0,0,0,....1...,0] ,you can also use MSE loss",
      "votes": null
    },
    {
      "id": "686317",
      "postDate": "12/03/2019 02:38:59",
      "content": "<p>I didn't get it. What do you mean by changing this to OH format?\nThe lengths of the documents are different, does this mean that the size of OH vector will be different?</p>",
      "rawMarkdown": "I didn't get it. What do you mean by changing this to OH format?\nThe lengths of the documents are different, does this mean that the size of OH vector will be different?",
      "votes": null
    },
    {
      "id": "686337",
      "postDate": "12/03/2019 03:26:55",
      "content": "<p>No,it should be padded to same length in a batch or simply in all data\nIn one-hot format,label[0,0,1] can compute MSE with logits[0.1,0.1,0.2] ,\nAnd ,simple start_positions label 2 should use CrossEntropyLoss.</p>",
      "rawMarkdown": "No,it should be padded to same length in a batch or simply in all data\nIn one-hot format,label[0,0,1] can compute MSE with logits[0.1,0.1,0.2] ,\nAnd ,simple start_positions label 2 should use CrossEntropyLoss.",
      "votes": null
    },
    {
      "id": "686524",
      "postDate": "12/03/2019 08:53:08",
      "content": "<p>Let's say, the doc length is 10 and the gold label for the answer is 6:9.</p>\n\n<p>You model tries to predict 2 vectors: pred probs for start indexes (<code>start_logits</code>), and same for end (<code>end_logits</code>).\nIf, say, \n<code>\nstart_logits = [0., 0., 0.9, 0., 0.1, 0., 0., 0., 0., 0.]\nend_logits   = [0., 0., 0.2, 0., 0., 0., 0., 0., 0.8, 0.]\n</code>\nThen your model will predict <code>argmax(start_logits):argmax(end_logits)</code> = 2:8.\nTherefore this specific loss function. </p>",
      "rawMarkdown": "Let's say, the doc length is 10 and the gold label for the answer is 6:9.\n\nYou model tries to predict 2 vectors: pred probs for start indexes (`start_logits`), and same for end (`end_logits`).\nIf, say, \n```\nstart_logits = [0., 0., 0.9, 0., 0.1, 0., 0., 0., 0., 0.]\nend_logits   = [0., 0., 0.2, 0., 0., 0., 0., 0., 0.8, 0.]\n```\nThen your model will predict `argmax(start_logits):argmax(end_logits)` = 2:8.\nTherefore this specific loss function.",
      "votes": null
    },
    {
      "id": "690707",
      "postDate": "12/09/2019 01:41:55",
      "content": "<p>I was understanding some outputs'shapes in the wrong way.\nThank you! </p>",
      "rawMarkdown": "I was understanding some outputs'shapes in the wrong way.\nThank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 686304,
      "author_name": "zhaomeng1126",
      "author_url": "",
      "post_date": "12/03/2019 02:17:57",
      "content": "<p>The most used label start_positions  maybe 5,10,100, if you change this to one-hot format[0,0,0,....1...,0] ,you can also use MSE loss</p>",
      "votes": null,
      "replies": [
        {
          "id": 686317,
          "author_name": "rinnqd",
          "author_url": "",
          "post_date": "12/03/2019 02:38:59",
          "content": "<p>I didn't get it. What do you mean by changing this to OH format?\nThe lengths of the documents are different, does this mean that the size of OH vector will be different?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686337,
          "author_name": "zhaomeng1126",
          "author_url": "",
          "post_date": "12/03/2019 03:26:55",
          "content": "<p>No,it should be padded to same length in a batch or simply in all data\nIn one-hot format,label[0,0,1] can compute MSE with logits[0.1,0.1,0.2] ,\nAnd ,simple start_positions label 2 should use CrossEntropyLoss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 686524,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "12/03/2019 08:53:08",
      "content": "<p>Let's say, the doc length is 10 and the gold label for the answer is 6:9.</p>\n\n<p>You model tries to predict 2 vectors: pred probs for start indexes (<code>start_logits</code>), and same for end (<code>end_logits</code>).\nIf, say, \n<code>\nstart_logits = [0., 0., 0.9, 0., 0.1, 0., 0., 0., 0., 0.]\nend_logits   = [0., 0., 0.2, 0., 0., 0., 0., 0., 0.8, 0.]\n</code>\nThen your model will predict <code>argmax(start_logits):argmax(end_logits)</code> = 2:8.\nTherefore this specific loss function. </p>",
      "votes": null,
      "replies": [
        {
          "id": 690707,
          "author_name": "rinnqd",
          "author_url": "",
          "post_date": "12/09/2019 01:41:55",
          "content": "<p>I was understanding some outputs'shapes in the wrong way.\nThank you! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "686246": "I have checked the model used in huggingface github for QuestionAnswering tasks and I  have found this ([link](https://github.com/huggingface/transformers/blob/master/transformers/modeling_bert.py))\n```\nloss_fct = CrossEntropyLoss(ignore_index=ignored_index)\n\nstart_loss = loss_fct(start_logits, start_positions)\n\nend_loss = loss_fct(end_logits, end_positions)\n\ntotal_loss = (start_loss + end_loss) / 2\n\n```\nIt is my first NLP competition and I  can't understand why the loss used here is CrossEntropyLoss. \nCan someone explain why we're using the CE loss instead of MSE?",
    "686304": "The most used label start_positions  maybe 5,10,100, if you change this to one-hot format[0,0,0,....1...,0] ,you can also use MSE loss",
    "686317": "I didn't get it. What do you mean by changing this to OH format?\nThe lengths of the documents are different, does this mean that the size of OH vector will be different?",
    "686337": "No,it should be padded to same length in a batch or simply in all data\nIn one-hot format,label[0,0,1] can compute MSE with logits[0.1,0.1,0.2] ,\nAnd ,simple start_positions label 2 should use CrossEntropyLoss.",
    "686524": "Let's say, the doc length is 10 and the gold label for the answer is 6:9.\n\nYou model tries to predict 2 vectors: pred probs for start indexes (`start_logits`), and same for end (`end_logits`).\nIf, say, \n```\nstart_logits = [0., 0., 0.9, 0., 0.1, 0., 0., 0., 0., 0.]\nend_logits   = [0., 0., 0.2, 0., 0., 0., 0., 0., 0.8, 0.]\n```\nThen your model will predict `argmax(start_logits):argmax(end_logits)` = 2:8.\nTherefore this specific loss function.",
    "690707": "I was understanding some outputs'shapes in the wrong way.\nThank you!"
  },
  "source": "meta"
}