{
  "id": 192245,
  "title": "Why is my loss function returning nan after a few steps?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/192245",
  "author_name": "",
  "post_date": "2020-10-20T18:02:01.967466500Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, I am trying to write code to create a TensorFlow model with a TPU (<code>torch_xla</code> was too memory inefficient, at least by my code).  I have converted Huan Vo's negative log loss function from this <a href=\"https://www.kaggle.com/huanvo/lyft-complete-train-and-prediction-pipeline\" target=\"_blank\">notebook</a>.  Here is what my loss function looks like:</p>\n<pre><code>def neg_multi_log_likelihood_batch(targets, target_availabilities, preds, confidences):\n    assert len(preds.shape) == 4, f\"Expected 4 dimensional array (B,M,T,C) but got a matrix with shape: {preds.shape}\"\n    batch_size, num_modes, future_len, num_coords = preds.shape\n\n    assert targets.shape == (batch_size, future_len, num_coords), f\"Expected 3 dimensional array of shape {(batch_size, future_len, num_coords)} but got a array of shape {targets.shape}\"\n    assert confidences.shape == (batch_size, num_modes), f\"Expected 2 dimensional array of shape {(batch_size, num_modes)} but got an array of shape {confidences.shape}\"\n\n    tf.expand_dims(targets, axis=1) # add the modes shape: (B, M, T, C)\n    target_availabilities = target_availabilities[:, None, :, None]  # add modes and cords (B, 1, T, 1) where Mode and coords are all none\n\n    # error (batch_size, num_modes, future_len)\n    error = K.sum(((targets - preds) * target_availabilities) ** 2, axis=-1) # reduce coords and use availability\n\n    with np.errstate(divide=\"ignore\"):  # when confidence is 0 log goes to -inf, but we're fine with it\n        # error (batch_size, num_modes)\n        error = tf.math.log(confidences) - 0.5 * K.sum(error, axis=-1)  # reduce time\n\n    # use max aggregator on modes for numerical stability\n    # error (batch_size, num_modes)\n    max_value = tf.math.reduce_max(error, axis=1, keepdims=True) # error are negative at this point, so max() gives the minimum one\n    error = -tf.math.log(K.sum(tf.math.exp(error - max_value), axis=-1, keepdims=True)) - max_value  # reduce modes\n    # print(\"error\", error)\n\n    return tf.math.reduce_mean(error)\n</code></pre>\n<p>It works fine for the first few steps, but then it just returns <code>nan</code> and the actual values don't come back.  What could be the problem?</p>",
  "messages": [
    {
      "id": "1055382",
      "postDate": "10/20/2020 18:02:01",
      "content": "<p>Hi, I am trying to write code to create a TensorFlow model with a TPU (<code>torch_xla</code> was too memory inefficient, at least by my code).  I have converted Huan Vo's negative log loss function from this <a href=\"https://www.kaggle.com/huanvo/lyft-complete-train-and-prediction-pipeline\" target=\"_blank\">notebook</a>.  Here is what my loss function looks like:</p>\n<pre><code>def neg_multi_log_likelihood_batch(targets, target_availabilities, preds, confidences):\n    assert len(preds.shape) == 4, f\"Expected 4 dimensional array (B,M,T,C) but got a matrix with shape: {preds.shape}\"\n    batch_size, num_modes, future_len, num_coords = preds.shape\n\n    assert targets.shape == (batch_size, future_len, num_coords), f\"Expected 3 dimensional array of shape {(batch_size, future_len, num_coords)} but got a array of shape {targets.shape}\"\n    assert confidences.shape == (batch_size, num_modes), f\"Expected 2 dimensional array of shape {(batch_size, num_modes)} but got an array of shape {confidences.shape}\"\n\n    tf.expand_dims(targets, axis=1) # add the modes shape: (B, M, T, C)\n    target_availabilities = target_availabilities[:, None, :, None]  # add modes and cords (B, 1, T, 1) where Mode and coords are all none\n\n    # error (batch_size, num_modes, future_len)\n    error = K.sum(((targets - preds) * target_availabilities) ** 2, axis=-1) # reduce coords and use availability\n\n    with np.errstate(divide=\"ignore\"):  # when confidence is 0 log goes to -inf, but we're fine with it\n        # error (batch_size, num_modes)\n        error = tf.math.log(confidences) - 0.5 * K.sum(error, axis=-1)  # reduce time\n\n    # use max aggregator on modes for numerical stability\n    # error (batch_size, num_modes)\n    max_value = tf.math.reduce_max(error, axis=1, keepdims=True) # error are negative at this point, so max() gives the minimum one\n    error = -tf.math.log(K.sum(tf.math.exp(error - max_value), axis=-1, keepdims=True)) - max_value  # reduce modes\n    # print(\"error\", error)\n\n    return tf.math.reduce_mean(error)\n</code></pre>\n<p>It works fine for the first few steps, but then it just returns <code>nan</code> and the actual values don't come back.  What could be the problem?</p>",
      "rawMarkdown": "Hi, I am trying to write code to create a TensorFlow model with a TPU (`torch_xla` was too memory inefficient, at least by my code).  I have converted Huan Vo's negative log loss function from this [notebook](https://www.kaggle.com/huanvo/lyft-complete-train-and-prediction-pipeline).  Here is what my loss function looks like:\n```python\ndef neg_multi_log_likelihood_batch(targets, target_availabilities, preds, confidences):\n    assert len(preds.shape) == 4, f\"Expected 4 dimensional array (B,M,T,C) but got a matrix with shape: {preds.shape}\"\n    batch_size, num_modes, future_len, num_coords = preds.shape\n    \n    assert targets.shape == (batch_size, future_len, num_coords), f\"Expected 3 dimensional array of shape {(batch_size, future_len, num_coords)} but got a array of shape {targets.shape}\"\n    assert confidences.shape == (batch_size, num_modes), f\"Expected 2 dimensional array of shape {(batch_size, num_modes)} but got an array of shape {confidences.shape}\"\n    \n    tf.expand_dims(targets, axis=1) # add the modes shape: (B, M, T, C)\n    target_availabilities = target_availabilities[:, None, :, None]  # add modes and cords (B, 1, T, 1) where Mode and coords are all none\n\n    # error (batch_size, num_modes, future_len)\n    error = K.sum(((targets - preds) * target_availabilities) ** 2, axis=-1) # reduce coords and use availability\n\n    with np.errstate(divide=\"ignore\"):  # when confidence is 0 log goes to -inf, but we're fine with it\n        # error (batch_size, num_modes)\n        error = tf.math.log(confidences) - 0.5 * K.sum(error, axis=-1)  # reduce time\n\n    # use max aggregator on modes for numerical stability\n    # error (batch_size, num_modes)\n    max_value = tf.math.reduce_max(error, axis=1, keepdims=True) # error are negative at this point, so max() gives the minimum one\n    error = -tf.math.log(K.sum(tf.math.exp(error - max_value), axis=-1, keepdims=True)) - max_value  # reduce modes\n    # print(\"error\", error)\n    \n    return tf.math.reduce_mean(error)\n```\nIt works fine for the first few steps, but then it just returns `nan` and the actual values don't come back.  What could be the problem?",
      "votes": null
    },
    {
      "id": "1055497",
      "postDate": "10/20/2020 21:23:58",
      "content": "<p>Please check out:</p>\n<p><a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773</a></p>",
      "rawMarkdown": "Please check out:\n\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773",
      "votes": null
    },
    {
      "id": "1055921",
      "postDate": "10/21/2020 08:26:37",
      "content": "<p>Check solution from here:  <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments#1049207\" target=\"_blank\">https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments#1049207</a></p>",
      "rawMarkdown": "Check solution from here:  https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments#1049207",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1055497,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "10/20/2020 21:23:58",
      "content": "<p>Please check out:</p>\n<p><a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1055921,
      "author_name": "loopdigga",
      "author_url": "",
      "post_date": "10/21/2020 08:26:37",
      "content": "<p>Check solution from here:  <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments#1049207\" target=\"_blank\">https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments#1049207</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1055382": "Hi, I am trying to write code to create a TensorFlow model with a TPU (`torch_xla` was too memory inefficient, at least by my code).  I have converted Huan Vo's negative log loss function from this [notebook](https://www.kaggle.com/huanvo/lyft-complete-train-and-prediction-pipeline).  Here is what my loss function looks like:\n```python\ndef neg_multi_log_likelihood_batch(targets, target_availabilities, preds, confidences):\n    assert len(preds.shape) == 4, f\"Expected 4 dimensional array (B,M,T,C) but got a matrix with shape: {preds.shape}\"\n    batch_size, num_modes, future_len, num_coords = preds.shape\n    \n    assert targets.shape == (batch_size, future_len, num_coords), f\"Expected 3 dimensional array of shape {(batch_size, future_len, num_coords)} but got a array of shape {targets.shape}\"\n    assert confidences.shape == (batch_size, num_modes), f\"Expected 2 dimensional array of shape {(batch_size, num_modes)} but got an array of shape {confidences.shape}\"\n    \n    tf.expand_dims(targets, axis=1) # add the modes shape: (B, M, T, C)\n    target_availabilities = target_availabilities[:, None, :, None]  # add modes and cords (B, 1, T, 1) where Mode and coords are all none\n\n    # error (batch_size, num_modes, future_len)\n    error = K.sum(((targets - preds) * target_availabilities) ** 2, axis=-1) # reduce coords and use availability\n\n    with np.errstate(divide=\"ignore\"):  # when confidence is 0 log goes to -inf, but we're fine with it\n        # error (batch_size, num_modes)\n        error = tf.math.log(confidences) - 0.5 * K.sum(error, axis=-1)  # reduce time\n\n    # use max aggregator on modes for numerical stability\n    # error (batch_size, num_modes)\n    max_value = tf.math.reduce_max(error, axis=1, keepdims=True) # error are negative at this point, so max() gives the minimum one\n    error = -tf.math.log(K.sum(tf.math.exp(error - max_value), axis=-1, keepdims=True)) - max_value  # reduce modes\n    # print(\"error\", error)\n    \n    return tf.math.reduce_mean(error)\n```\nIt works fine for the first few steps, but then it just returns `nan` and the actual values don't come back.  What could be the problem?",
    "1055497": "Please check out:\n\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773",
    "1055921": "Check solution from here:  https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments#1049207"
  },
  "source": "meta"
}