{
  "id": 556193,
  "title": "NN returning no loss scoring.",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556193",
  "author_name": "",
  "post_date": "2025-01-11T17:48:59.034351600Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi everyone! I've been experimenting with using symbol_id in an embedding layer in a tensorflow NN and have borrowed a structure from this notebook <a href=\"https://www.kaggle.com/code/colinmorris/embedding-layers#Good-idea:-Embedding-layers\" target=\"_blank\">https://www.kaggle.com/code/colinmorris/embedding-layers#Good-idea:-Embedding-layers</a>.<br>\nHowever, when training I always get a nan train/val loss. I'm certain that there's something wrong with my model and I've missed something silly since I'm relatively new to deep learning but I cannot seem to spot it at the moment. I've included my model code below, any help would be much appreciated!</p>\n<p>`hidden_units = (128,64,32) #basic pyramid shape for initial attempt<br>\nstock_embedding_size = 8 #dimensionality of output</p>\n<p>cat_data = X_train['symbol_id']</p>\n<p>def short_term_model():</p>\n<pre><code>\nsymbol_id_input = keras.Input(shape=(1,), =)\nnum_input = keras.Input(shape=(len(CONFIG.feature_cols + CONFIG.lag_cols + CONFIG.timeseries_cols),), =)\n\n\n\nsymbol_embedded = keras.layers.Embedding(max(cat_data)+1, stock_embedding_size, \n                                       =1, =)(symbol_id_input)\nsymbol_flattened = keras.layers.Flatten()(symbol_embedded)\nout = keras.layers.Concatenate()([symbol_flattened, num_input])\n\n\n n_hidden  hidden_units:\n\n    #out = keras.layers.BatchNormalization()(out)\n    out = keras.layers.Dense(n_hidden, =)(out)\n    #out = keras.layers.Dropout(=0.02,seed=42)(out)\n\n\n\n\nout = keras.layers.Dense(1, =, =)(out)\n\nmodel = keras.Model(\ninputs = [symbol_id_input, num_input],\noutputs = out,\n)\n\nreturn model`\n</code></pre>",
  "messages": [
    {
      "id": "3094165",
      "postDate": "01/11/2025 17:48:59",
      "content": "<p>Hi everyone! I've been experimenting with using symbol_id in an embedding layer in a tensorflow NN and have borrowed a structure from this notebook <a href=\"https://www.kaggle.com/code/colinmorris/embedding-layers#Good-idea:-Embedding-layers\" target=\"_blank\">https://www.kaggle.com/code/colinmorris/embedding-layers#Good-idea:-Embedding-layers</a>.<br>\nHowever, when training I always get a nan train/val loss. I'm certain that there's something wrong with my model and I've missed something silly since I'm relatively new to deep learning but I cannot seem to spot it at the moment. I've included my model code below, any help would be much appreciated!</p>\n<p>`hidden_units = (128,64,32) #basic pyramid shape for initial attempt<br>\nstock_embedding_size = 8 #dimensionality of output</p>\n<p>cat_data = X_train['symbol_id']</p>\n<p>def short_term_model():</p>\n<pre><code>\nsymbol_id_input = keras.Input(shape=(1,), =)\nnum_input = keras.Input(shape=(len(CONFIG.feature_cols + CONFIG.lag_cols + CONFIG.timeseries_cols),), =)\n\n\n\nsymbol_embedded = keras.layers.Embedding(max(cat_data)+1, stock_embedding_size, \n                                       =1, =)(symbol_id_input)\nsymbol_flattened = keras.layers.Flatten()(symbol_embedded)\nout = keras.layers.Concatenate()([symbol_flattened, num_input])\n\n\n n_hidden  hidden_units:\n\n    #out = keras.layers.BatchNormalization()(out)\n    out = keras.layers.Dense(n_hidden, =)(out)\n    #out = keras.layers.Dropout(=0.02,seed=42)(out)\n\n\n\n\nout = keras.layers.Dense(1, =, =)(out)\n\nmodel = keras.Model(\ninputs = [symbol_id_input, num_input],\noutputs = out,\n)\n\nreturn model`\n</code></pre>",
      "rawMarkdown": "Hi everyone! I've been experimenting with using symbol_id in an embedding layer in a tensorflow NN and have borrowed a structure from this notebook https://www.kaggle.com/code/colinmorris/embedding-layers#Good-idea:-Embedding-layers.\nHowever, when training I always get a nan train/val loss. I'm certain that there's something wrong with my model and I've missed something silly since I'm relatively new to deep learning but I cannot seem to spot it at the moment. I've included my model code below, any help would be much appreciated!\n\n`hidden_units = (128,64,32) #basic pyramid shape for initial attempt\nstock_embedding_size = 8 #dimensionality of output\n\ncat_data = X_train['symbol_id']\n\ndef short_term_model():\n    \n    # Splitting the inputs to the symbol id for embedding and the rest of the numerical data\n    symbol_id_input = keras.Input(shape=(1,), name='symbol_id')\n    num_input = keras.Input(shape=(len(CONFIG.feature_cols + CONFIG.lag_cols + CONFIG.timeseries_cols),), name='num_data')\n\n\n    #embedding, flatenning and concatenating\n    symbol_embedded = keras.layers.Embedding(max(cat_data)+1, stock_embedding_size, \n                                           input_length=1, name='symbol_embedding')(symbol_id_input)\n    symbol_flattened = keras.layers.Flatten()(symbol_embedded)\n    out = keras.layers.Concatenate()([symbol_flattened, num_input])\n    \n    # Add one or more hidden layers\n    for n_hidden in hidden_units:\n\n        #out = keras.layers.BatchNormalization()(out)\n        out = keras.layers.Dense(n_hidden, activation='swish')(out)\n        #out = keras.layers.Dropout(rate=0.02,seed=42)(out)\n        \n\n\n    # A single output: our predicted rating\n    out = keras.layers.Dense(1, activation='linear', name='prediction')(out)\n    \n    model = keras.Model(\n    inputs = [symbol_id_input, num_input],\n    outputs = out,\n    )\n    \n    return model`",
      "votes": null
    },
    {
      "id": "3094170",
      "postDate": "01/11/2025 17:52:25",
      "content": "<p>did you fill missing values in your input X ?</p>",
      "rawMarkdown": "did you fill missing values in your input X ?",
      "votes": null
    },
    {
      "id": "3094171",
      "postDate": "01/11/2025 17:52:25",
      "content": "<p>do you fill nans?</p>",
      "rawMarkdown": "do you fill nans?",
      "votes": null
    },
    {
      "id": "3094182",
      "postDate": "01/11/2025 18:00:48",
      "content": "<p>Hi, yes I have imputed any missing values. I'm assuming if I got my input shapes incorrect when feeding for training I'd get an error message?</p>",
      "rawMarkdown": "Hi, yes I have imputed any missing values. I'm assuming if I got my input shapes incorrect when feeding for training I'd get an error message?",
      "votes": null
    },
    {
      "id": "3094269",
      "postDate": "01/11/2025 20:31:26",
      "content": "<p>I see you've already been asked about filling missing values, but what about scaling - did you use some?</p>",
      "rawMarkdown": "I see you've already been asked about filling missing values, but what about scaling - did you use some?",
      "votes": null
    },
    {
      "id": "3094294",
      "postDate": "01/11/2025 21:50:15",
      "content": "<p>What error are you using? </p>",
      "rawMarkdown": "What error are you using?",
      "votes": null
    },
    {
      "id": "3094298",
      "postDate": "01/11/2025 22:03:02",
      "content": "<p>Hello, did the loss start as NaN right away? If so, there might be an issue with your dataset. Check your data loader (if you're using one), verify for NaNs, and ensure your tensor shapes are correct. If the loss didn’t start as NaN but later exploded, the problem could lie in your loss function. For example, if you're using R² as a loss, it might lead to instability. Instead, consider using 1 - R², which tends to stabilize the training process.</p>",
      "rawMarkdown": "Hello, did the loss start as NaN right away? If so, there might be an issue with your dataset. Check your data loader (if you're using one), verify for NaNs, and ensure your tensor shapes are correct. If the loss didn’t start as NaN but later exploded, the problem could lie in your loss function. For example, if you're using R² as a loss, it might lead to instability. Instead, consider using 1 - R², which tends to stabilize the training process.",
      "votes": null
    },
    {
      "id": "3094691",
      "postDate": "01/12/2025 12:19:28",
      "content": "<p>Hi, error started off as NaN right away. I thought it could be that I was feeding polars sets to the model (which I have learned is not supported) but after a pandas conversion it's still NaN. I'm going to take a look at the shapes of my inputs. On the topic of polars, how would you reccommend converting a polars df into an object that tensorflow can accept which doesn't blow up memory usage? I've tried to numpy and to pandas with little success so far. Thanks.</p>",
      "rawMarkdown": "Hi, error started off as NaN right away. I thought it could be that I was feeding polars sets to the model (which I have learned is not supported) but after a pandas conversion it's still NaN. I'm going to take a look at the shapes of my inputs. On the topic of polars, how would you reccommend converting a polars df into an object that tensorflow can accept which doesn't blow up memory usage? I've tried to numpy and to pandas with little success so far. Thanks.",
      "votes": null
    },
    {
      "id": "3094846",
      "postDate": "01/12/2025 15:52:49",
      "content": "<p>Hi all, found the issue causing this and it was a rather silly mistake on my part. Thanks for all the help!</p>",
      "rawMarkdown": "Hi all, found the issue causing this and it was a rather silly mistake on my part. Thanks for all the help!",
      "votes": null
    },
    {
      "id": "3094849",
      "postDate": "01/12/2025 15:54:48",
      "content": "<p>TensorFlow supports both Pandas DataFrames and NumPy arrays, and converting from Polars to these formats is straightforward. To facilitate debugging, consider using a simpler architecture and a smaller subset of data. If you're using a custom loss function, try starting with a basic one, such as MSE.</p>",
      "rawMarkdown": "TensorFlow supports both Pandas DataFrames and NumPy arrays, and converting from Polars to these formats is straightforward. To facilitate debugging, consider using a simpler architecture and a smaller subset of data. If you're using a custom loss function, try starting with a basic one, such as MSE.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3094170,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "01/11/2025 17:52:25",
      "content": "<p>did you fill missing values in your input X ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3094182,
          "author_name": "jacobhh1",
          "author_url": "",
          "post_date": "01/11/2025 18:00:48",
          "content": "<p>Hi, yes I have imputed any missing values. I'm assuming if I got my input shapes incorrect when feeding for training I'd get an error message?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3094171,
      "author_name": "sritichaimae",
      "author_url": "",
      "post_date": "01/11/2025 17:52:25",
      "content": "<p>do you fill nans?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3094269,
      "author_name": "yekenot",
      "author_url": "",
      "post_date": "01/11/2025 20:31:26",
      "content": "<p>I see you've already been asked about filling missing values, but what about scaling - did you use some?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3094294,
      "author_name": "redfoongus",
      "author_url": "",
      "post_date": "01/11/2025 21:50:15",
      "content": "<p>What error are you using? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3094298,
      "author_name": "moujahidboulaouz",
      "author_url": "",
      "post_date": "01/11/2025 22:03:02",
      "content": "<p>Hello, did the loss start as NaN right away? If so, there might be an issue with your dataset. Check your data loader (if you're using one), verify for NaNs, and ensure your tensor shapes are correct. If the loss didn’t start as NaN but later exploded, the problem could lie in your loss function. For example, if you're using R² as a loss, it might lead to instability. Instead, consider using 1 - R², which tends to stabilize the training process.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3094691,
          "author_name": "jacobhh1",
          "author_url": "",
          "post_date": "01/12/2025 12:19:28",
          "content": "<p>Hi, error started off as NaN right away. I thought it could be that I was feeding polars sets to the model (which I have learned is not supported) but after a pandas conversion it's still NaN. I'm going to take a look at the shapes of my inputs. On the topic of polars, how would you reccommend converting a polars df into an object that tensorflow can accept which doesn't blow up memory usage? I've tried to numpy and to pandas with little success so far. Thanks.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3094849,
              "author_name": "moujahidboulaouz",
              "author_url": "",
              "post_date": "01/12/2025 15:54:48",
              "content": "<p>TensorFlow supports both Pandas DataFrames and NumPy arrays, and converting from Polars to these formats is straightforward. To facilitate debugging, consider using a simpler architecture and a smaller subset of data. If you're using a custom loss function, try starting with a basic one, such as MSE.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3094846,
      "author_name": "jacobhh1",
      "author_url": "",
      "post_date": "01/12/2025 15:52:49",
      "content": "<p>Hi all, found the issue causing this and it was a rather silly mistake on my part. Thanks for all the help!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3094165": "Hi everyone! I've been experimenting with using symbol_id in an embedding layer in a tensorflow NN and have borrowed a structure from this notebook https://www.kaggle.com/code/colinmorris/embedding-layers#Good-idea:-Embedding-layers.\nHowever, when training I always get a nan train/val loss. I'm certain that there's something wrong with my model and I've missed something silly since I'm relatively new to deep learning but I cannot seem to spot it at the moment. I've included my model code below, any help would be much appreciated!\n\n`hidden_units = (128,64,32) #basic pyramid shape for initial attempt\nstock_embedding_size = 8 #dimensionality of output\n\ncat_data = X_train['symbol_id']\n\ndef short_term_model():\n    \n    # Splitting the inputs to the symbol id for embedding and the rest of the numerical data\n    symbol_id_input = keras.Input(shape=(1,), name='symbol_id')\n    num_input = keras.Input(shape=(len(CONFIG.feature_cols + CONFIG.lag_cols + CONFIG.timeseries_cols),), name='num_data')\n\n\n    #embedding, flatenning and concatenating\n    symbol_embedded = keras.layers.Embedding(max(cat_data)+1, stock_embedding_size, \n                                           input_length=1, name='symbol_embedding')(symbol_id_input)\n    symbol_flattened = keras.layers.Flatten()(symbol_embedded)\n    out = keras.layers.Concatenate()([symbol_flattened, num_input])\n    \n    # Add one or more hidden layers\n    for n_hidden in hidden_units:\n\n        #out = keras.layers.BatchNormalization()(out)\n        out = keras.layers.Dense(n_hidden, activation='swish')(out)\n        #out = keras.layers.Dropout(rate=0.02,seed=42)(out)\n        \n\n\n    # A single output: our predicted rating\n    out = keras.layers.Dense(1, activation='linear', name='prediction')(out)\n    \n    model = keras.Model(\n    inputs = [symbol_id_input, num_input],\n    outputs = out,\n    )\n    \n    return model`",
    "3094170": "did you fill missing values in your input X ?",
    "3094171": "do you fill nans?",
    "3094182": "Hi, yes I have imputed any missing values. I'm assuming if I got my input shapes incorrect when feeding for training I'd get an error message?",
    "3094269": "I see you've already been asked about filling missing values, but what about scaling - did you use some?",
    "3094294": "What error are you using?",
    "3094298": "Hello, did the loss start as NaN right away? If so, there might be an issue with your dataset. Check your data loader (if you're using one), verify for NaNs, and ensure your tensor shapes are correct. If the loss didn’t start as NaN but later exploded, the problem could lie in your loss function. For example, if you're using R² as a loss, it might lead to instability. Instead, consider using 1 - R², which tends to stabilize the training process.",
    "3094691": "Hi, error started off as NaN right away. I thought it could be that I was feeding polars sets to the model (which I have learned is not supported) but after a pandas conversion it's still NaN. I'm going to take a look at the shapes of my inputs. On the topic of polars, how would you reccommend converting a polars df into an object that tensorflow can accept which doesn't blow up memory usage? I've tried to numpy and to pandas with little success so far. Thanks.",
    "3094846": "Hi all, found the issue causing this and it was a rather silly mistake on my part. Thanks for all the help!",
    "3094849": "TensorFlow supports both Pandas DataFrames and NumPy arrays, and converting from Polars to these formats is straightforward. To facilitate debugging, consider using a simpler architecture and a smaller subset of data. If you're using a custom loss function, try starting with a basic one, such as MSE."
  },
  "source": "meta"
}