{
  "id": 113928,
  "title": "Multiple output and multiple loss functions in Keras",
  "url": "/competitions/pku-autonomous-driving/discussion/113928",
  "author_name": "",
  "post_date": "2019-10-23T04:13:00.140920300Z",
  "votes": 22,
  "comment_count": 6,
  "views": 0,
  "content": "<p>One thing you might have noticed is that this competition requires two different types of losses: cross-entropy (for binary confidence that a certain brand of car is in the image), and MSE (for every other output target). This means that you might need to not only use two outputs in your model, but that each of those two outputs will have a different loss function. For those who are less familiar with Keras, I highly recommend to read this excellent guide about the functional API: <a href=\"https://keras.io/getting-started/functional-api-guide/\">https://keras.io/getting-started/functional-api-guide/</a></p>\n\n<p>I wanted to specifically point out this part:\n```python\nmain_input = Input(shape=(100,), dtype='int32', name='main_input')\n...\nlstm_out = LSTM(32)(x)</p>\n\n<p>auxiliary_output = Dense(1, activation='sigmoid', name='aux_output')(lstm_out)</p>\n\n<p>auxiliary_input = Input(shape=(5,), name='aux_input')\nx = keras.layers.concatenate([lstm_out, auxiliary_input])\n...\nmain_output = Dense(1, activation='sigmoid', name='main_output')(x)</p>\n\n<p>model = Model(inputs=[main_input, auxiliary_input], outputs=[main_output, auxiliary_output])\n```\nAs you can see, you can have two different outputs: the main output (e.g. confidence level, which is between 0 and 1) and the aux output (yaw, pitch, roll, x, y, z), where each could have a different activation function (respectively, sigmoid and linear). However, even more interesting is this part:</p>\n\n<p>```python\nmodel.compile(optimizer='rmsprop',\n              loss={'main_output': 'binary_crossentropy', 'aux_output': 'binary_crossentropy'},\n              loss_weights={'main_output': 1., 'aux_output': 0.2})</p>\n\n<h1>And trained it via:</h1>\n\n<p>model.fit({'main_input': headline_data, 'aux_input': additional_data},\n          {'main_output': headline_labels, 'aux_output': additional_labels},\n          epochs=50, batch_size=32)\n```\nas you can see, you can pass a dictionary to the loss parameter, which contains the loss you want to specify for each of the named output (make sure you gave a name when you created the output layer function!). This means you can specify one of the outputs to have a BCE loss (e.g. for confidence), and the other to have a MSE loss (since it's an exact value you are trying to predict).</p>\n\n<p>Hope this short write-up was useful. Let me know your thoughts :)</p>",
  "messages": [
    {
      "id": "655448",
      "postDate": "10/23/2019 04:13:00",
      "content": "<p>One thing you might have noticed is that this competition requires two different types of losses: cross-entropy (for binary confidence that a certain brand of car is in the image), and MSE (for every other output target). This means that you might need to not only use two outputs in your model, but that each of those two outputs will have a different loss function. For those who are less familiar with Keras, I highly recommend to read this excellent guide about the functional API: <a href=\"https://keras.io/getting-started/functional-api-guide/\">https://keras.io/getting-started/functional-api-guide/</a></p>\n\n<p>I wanted to specifically point out this part:\n```python\nmain_input = Input(shape=(100,), dtype='int32', name='main_input')\n...\nlstm_out = LSTM(32)(x)</p>\n\n<p>auxiliary_output = Dense(1, activation='sigmoid', name='aux_output')(lstm_out)</p>\n\n<p>auxiliary_input = Input(shape=(5,), name='aux_input')\nx = keras.layers.concatenate([lstm_out, auxiliary_input])\n...\nmain_output = Dense(1, activation='sigmoid', name='main_output')(x)</p>\n\n<p>model = Model(inputs=[main_input, auxiliary_input], outputs=[main_output, auxiliary_output])\n```\nAs you can see, you can have two different outputs: the main output (e.g. confidence level, which is between 0 and 1) and the aux output (yaw, pitch, roll, x, y, z), where each could have a different activation function (respectively, sigmoid and linear). However, even more interesting is this part:</p>\n\n<p>```python\nmodel.compile(optimizer='rmsprop',\n              loss={'main_output': 'binary_crossentropy', 'aux_output': 'binary_crossentropy'},\n              loss_weights={'main_output': 1., 'aux_output': 0.2})</p>\n\n<h1>And trained it via:</h1>\n\n<p>model.fit({'main_input': headline_data, 'aux_input': additional_data},\n          {'main_output': headline_labels, 'aux_output': additional_labels},\n          epochs=50, batch_size=32)\n```\nas you can see, you can pass a dictionary to the loss parameter, which contains the loss you want to specify for each of the named output (make sure you gave a name when you created the output layer function!). This means you can specify one of the outputs to have a BCE loss (e.g. for confidence), and the other to have a MSE loss (since it's an exact value you are trying to predict).</p>\n\n<p>Hope this short write-up was useful. Let me know your thoughts :)</p>",
      "rawMarkdown": "One thing you might have noticed is that this competition requires two different types of losses: cross-entropy (for binary confidence that a certain brand of car is in the image), and MSE (for every other output target). This means that you might need to not only use two outputs in your model, but that each of those two outputs will have a different loss function. For those who are less familiar with Keras, I highly recommend to read this excellent guide about the functional API: https://keras.io/getting-started/functional-api-guide/\n\nI wanted to specifically point out this part:\n```python\nmain_input = Input(shape=(100,), dtype='int32', name='main_input')\n...\nlstm_out = LSTM(32)(x)\n\nauxiliary_output = Dense(1, activation='sigmoid', name='aux_output')(lstm_out)\n\nauxiliary_input = Input(shape=(5,), name='aux_input')\nx = keras.layers.concatenate([lstm_out, auxiliary_input])\n...\nmain_output = Dense(1, activation='sigmoid', name='main_output')(x)\n\nmodel = Model(inputs=[main_input, auxiliary_input], outputs=[main_output, auxiliary_output])\n```\nAs you can see, you can have two different outputs: the main output (e.g. confidence level, which is between 0 and 1) and the aux output (yaw, pitch, roll, x, y, z), where each could have a different activation function (respectively, sigmoid and linear). However, even more interesting is this part:\n\n```python\nmodel.compile(optimizer='rmsprop',\n              loss={'main_output': 'binary_crossentropy', 'aux_output': 'binary_crossentropy'},\n              loss_weights={'main_output': 1., 'aux_output': 0.2})\n\n# And trained it via:\nmodel.fit({'main_input': headline_data, 'aux_input': additional_data},\n          {'main_output': headline_labels, 'aux_output': additional_labels},\n          epochs=50, batch_size=32)\n```\nas you can see, you can pass a dictionary to the loss parameter, which contains the loss you want to specify for each of the named output (make sure you gave a name when you created the output layer function!). This means you can specify one of the outputs to have a BCE loss (e.g. for confidence), and the other to have a MSE loss (since it's an exact value you are trying to predict).\n\nHope this short write-up was useful. Let me know your thoughts :)",
      "votes": null
    },
    {
      "id": "655911",
      "postDate": "10/23/2019 16:58:05",
      "content": "<p>With this approach does it matter what the loss weights are? I looked it up and people say that BCE should have a low weight and MSE high.</p>",
      "rawMarkdown": "With this approach does it matter what the loss weights are? I looked it up and people say that BCE should have a low weight and MSE high.",
      "votes": null
    },
    {
      "id": "656132",
      "postDate": "10/24/2019 00:08:13",
      "content": "<p>This is one more hyperparameter you can tune ;)</p>",
      "rawMarkdown": "This is one more hyperparameter you can tune ;)",
      "votes": null
    },
    {
      "id": "656210",
      "postDate": "10/24/2019 02:34:36",
      "content": "<p>good point.</p>",
      "rawMarkdown": "good point.",
      "votes": null
    },
    {
      "id": "656686",
      "postDate": "10/24/2019 14:31:09",
      "content": "<p>Thanks for sharing <a href=\"/xhlulu\">@xhlulu</a> !</p>",
      "rawMarkdown": "Thanks for sharing @xhlulu !",
      "votes": null
    },
    {
      "id": "685848",
      "postDate": "12/02/2019 12:05:11",
      "content": "<p>Good job!\nI see others' code, the last layer is a Lambda layer, so is this one output or two?</p>",
      "rawMarkdown": "Good job!\nI see others' code, the last layer is a Lambda layer, so is this one output or two?",
      "votes": null
    },
    {
      "id": "814420",
      "postDate": "04/20/2020 16:43:26",
      "content": "<p>what if I did separate function to calculate the total loss, how can I make x_output weighted twice y_output?</p>\n\n<p>model.compile(optimizer='rmsprop',\n              loss= loss_function,\n              loss_weights={'main_output': 1., 'aux_output': 0.2})</p>\n\n<p>def loss_fucntion (targets, model_output):\nx_target = labels['x_target']\ny_target = labels['y_target']\nmain_output, aux_output = model_outputs</p>\n\n<pre><code>x_loss = tf.keras.backend.sparse_categorical_crossentropy\n    x_target, main_output, from_logits=True)\n\ny_loss = tf.keras.backend.sparse_categorical_crossentropy(\n    y_target, aux_output, from_logits=True)\n\ntotal_loss = (tf.reduce_mean(x_loss) + tf.reduce_mean(y_loss)) / 2\n\nreturn total_loss\n</code></pre>",
      "rawMarkdown": "what if I did separate function to calculate the total loss, how can I make x_output weighted twice y_output?\n\nmodel.compile(optimizer='rmsprop',\n              loss= loss_function,\n              loss_weights={'main_output': 1., 'aux_output': 0.2})\n\ndef loss_fucntion (targets, model_output):\nx_target = labels['x_target']\ny_target = labels['y_target']\nmain_output, aux_output = model_outputs\n\n    x_loss = tf.keras.backend.sparse_categorical_crossentropy\n        x_target, main_output, from_logits=True)\n    \n    y_loss = tf.keras.backend.sparse_categorical_crossentropy(\n        y_target, aux_output, from_logits=True)\n    \n    total_loss = (tf.reduce_mean(x_loss) + tf.reduce_mean(y_loss)) / 2\n\n    return total_loss",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 655911,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "10/23/2019 16:58:05",
      "content": "<p>With this approach does it matter what the loss weights are? I looked it up and people say that BCE should have a low weight and MSE high.</p>",
      "votes": null,
      "replies": [
        {
          "id": 656132,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "10/24/2019 00:08:13",
          "content": "<p>This is one more hyperparameter you can tune ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 656210,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "10/24/2019 02:34:36",
      "content": "<p>good point.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 656686,
      "author_name": "phunghieu",
      "author_url": "",
      "post_date": "10/24/2019 14:31:09",
      "content": "<p>Thanks for sharing <a href=\"/xhlulu\">@xhlulu</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 685848,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/02/2019 12:05:11",
      "content": "<p>Good job!\nI see others' code, the last layer is a Lambda layer, so is this one output or two?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 814420,
      "author_name": "sehamra",
      "author_url": "",
      "post_date": "04/20/2020 16:43:26",
      "content": "<p>what if I did separate function to calculate the total loss, how can I make x_output weighted twice y_output?</p>\n\n<p>model.compile(optimizer='rmsprop',\n              loss= loss_function,\n              loss_weights={'main_output': 1., 'aux_output': 0.2})</p>\n\n<p>def loss_fucntion (targets, model_output):\nx_target = labels['x_target']\ny_target = labels['y_target']\nmain_output, aux_output = model_outputs</p>\n\n<pre><code>x_loss = tf.keras.backend.sparse_categorical_crossentropy\n    x_target, main_output, from_logits=True)\n\ny_loss = tf.keras.backend.sparse_categorical_crossentropy(\n    y_target, aux_output, from_logits=True)\n\ntotal_loss = (tf.reduce_mean(x_loss) + tf.reduce_mean(y_loss)) / 2\n\nreturn total_loss\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "655448": "One thing you might have noticed is that this competition requires two different types of losses: cross-entropy (for binary confidence that a certain brand of car is in the image), and MSE (for every other output target). This means that you might need to not only use two outputs in your model, but that each of those two outputs will have a different loss function. For those who are less familiar with Keras, I highly recommend to read this excellent guide about the functional API: https://keras.io/getting-started/functional-api-guide/\n\nI wanted to specifically point out this part:\n```python\nmain_input = Input(shape=(100,), dtype='int32', name='main_input')\n...\nlstm_out = LSTM(32)(x)\n\nauxiliary_output = Dense(1, activation='sigmoid', name='aux_output')(lstm_out)\n\nauxiliary_input = Input(shape=(5,), name='aux_input')\nx = keras.layers.concatenate([lstm_out, auxiliary_input])\n...\nmain_output = Dense(1, activation='sigmoid', name='main_output')(x)\n\nmodel = Model(inputs=[main_input, auxiliary_input], outputs=[main_output, auxiliary_output])\n```\nAs you can see, you can have two different outputs: the main output (e.g. confidence level, which is between 0 and 1) and the aux output (yaw, pitch, roll, x, y, z), where each could have a different activation function (respectively, sigmoid and linear). However, even more interesting is this part:\n\n```python\nmodel.compile(optimizer='rmsprop',\n              loss={'main_output': 'binary_crossentropy', 'aux_output': 'binary_crossentropy'},\n              loss_weights={'main_output': 1., 'aux_output': 0.2})\n\n# And trained it via:\nmodel.fit({'main_input': headline_data, 'aux_input': additional_data},\n          {'main_output': headline_labels, 'aux_output': additional_labels},\n          epochs=50, batch_size=32)\n```\nas you can see, you can pass a dictionary to the loss parameter, which contains the loss you want to specify for each of the named output (make sure you gave a name when you created the output layer function!). This means you can specify one of the outputs to have a BCE loss (e.g. for confidence), and the other to have a MSE loss (since it's an exact value you are trying to predict).\n\nHope this short write-up was useful. Let me know your thoughts :)",
    "655911": "With this approach does it matter what the loss weights are? I looked it up and people say that BCE should have a low weight and MSE high.",
    "656132": "This is one more hyperparameter you can tune ;)",
    "656210": "good point.",
    "656686": "Thanks for sharing @xhlulu !",
    "685848": "Good job!\nI see others' code, the last layer is a Lambda layer, so is this one output or two?",
    "814420": "what if I did separate function to calculate the total loss, how can I make x_output weighted twice y_output?\n\nmodel.compile(optimizer='rmsprop',\n              loss= loss_function,\n              loss_weights={'main_output': 1., 'aux_output': 0.2})\n\ndef loss_fucntion (targets, model_output):\nx_target = labels['x_target']\ny_target = labels['y_target']\nmain_output, aux_output = model_outputs\n\n    x_loss = tf.keras.backend.sparse_categorical_crossentropy\n        x_target, main_output, from_logits=True)\n    \n    y_loss = tf.keras.backend.sparse_categorical_crossentropy(\n        y_target, aux_output, from_logits=True)\n    \n    total_loss = (tf.reduce_mean(x_loss) + tf.reduce_mean(y_loss)) / 2\n\n    return total_loss"
  },
  "source": "meta"
}