{
  "id": 550453,
  "title": "Do anyone have any insights / general tips about using neural networks?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/550453",
  "author_name": "",
  "post_date": "2024-12-07T15:23:24.209494900Z",
  "votes": 11,
  "comment_count": 11,
  "views": 0,
  "content": "<p>So far, all my attempts at using neural networks have yielded far worse performance than any GBDT based model, or even linear regression. It ends up generalizing catastrophically to my validation.</p>\n<p>If I use something like lasso regression, I get a validation score of like 0.0035. If I use neural networks I get validation scores -10.</p>\n<p>I use a basic NN architecture. Dense, feedforward, GELU, dropout, layernorm, tuned with Adam. I've tried to do hyperparameter search over basically all the parameters (learning rate, dropout rate, l2 norm, momentum parameters, number of layers, dimension of hidden layers).</p>\n<p>I'm wondering if I'm doing too few gradient steps per batch? I'm just going through the dataset chronologically one time. Should I be doing more than one gradient step per sample in the dataset?</p>",
  "messages": [
    {
      "id": "3066038",
      "postDate": "12/07/2024 15:23:24",
      "content": "<p>So far, all my attempts at using neural networks have yielded far worse performance than any GBDT based model, or even linear regression. It ends up generalizing catastrophically to my validation.</p>\n<p>If I use something like lasso regression, I get a validation score of like 0.0035. If I use neural networks I get validation scores -10.</p>\n<p>I use a basic NN architecture. Dense, feedforward, GELU, dropout, layernorm, tuned with Adam. I've tried to do hyperparameter search over basically all the parameters (learning rate, dropout rate, l2 norm, momentum parameters, number of layers, dimension of hidden layers).</p>\n<p>I'm wondering if I'm doing too few gradient steps per batch? I'm just going through the dataset chronologically one time. Should I be doing more than one gradient step per sample in the dataset?</p>",
      "rawMarkdown": "So far, all my attempts at using neural networks have yielded far worse performance than any GBDT based model, or even linear regression. It ends up generalizing catastrophically to my validation.\n\nIf I use something like lasso regression, I get a validation score of like 0.0035. If I use neural networks I get validation scores -10.\n\nI use a basic NN architecture. Dense, feedforward, GELU, dropout, layernorm, tuned with Adam. I've tried to do hyperparameter search over basically all the parameters (learning rate, dropout rate, l2 norm, momentum parameters, number of layers, dimension of hidden layers).\n\nI'm wondering if I'm doing too few gradient steps per batch? I'm just going through the dataset chronologically one time. Should I be doing more than one gradient step per sample in the dataset?",
      "votes": null
    },
    {
      "id": "3066248",
      "postDate": "12/07/2024 21:42:05",
      "content": "<blockquote>\n  <p>I'm just going through the dataset chronologically one time</p>\n</blockquote>\n<p>No, you should shuffle it.</p>\n<p>small lr, simple articheture (do not go deep), weighted r2 score as loss, you should got improvement coz my MLP got 0.0050 score on online test. Let me know hen you got any improvement of result.</p>",
      "rawMarkdown": "> I'm just going through the dataset chronologically one time\n\nNo, you should shuffle it.\n\nsmall lr, simple articheture (do not go deep), weighted r2 score as loss, you should got improvement coz my MLP got 0.0050 score on online test. Let me know hen you got any improvement of result.",
      "votes": null
    },
    {
      "id": "3068328",
      "postDate": "12/10/2024 06:43:16",
      "content": "<p>Something seems to be off, you should be able to get a decent score with a NN model.  What kind of score do you get on a validation portion of the training data -- say the last 120 days?  You'd probably have to share a bit more about what you're doing in order for anyone to give you specific pointers.  There is a public notebook that includes a basic fully connected NN, you should probably check that one out and see whether it produces different results from what you're getting.  I don't think you need anything too fancy on the NN model to get performance in the lasso range IMHO.</p>",
      "rawMarkdown": "Something seems to be off, you should be able to get a decent score with a NN model.  What kind of score do you get on a validation portion of the training data -- say the last 120 days?  You'd probably have to share a bit more about what you're doing in order for anyone to give you specific pointers.  There is a public notebook that includes a basic fully connected NN, you should probably check that one out and see whether it produces different results from what you're getting.  I don't think you need anything too fancy on the NN model to get performance in the lasso range IMHO.",
      "votes": null
    },
    {
      "id": "3069422",
      "postDate": "12/11/2024 13:26:19",
      "content": "<p>Some things that helped me.</p>\n<ol>\n<li>remove weights from loss_fn during training</li>\n<li>remove dropout</li>\n<li>remove layer norm</li>\n<li>remove gradient clipping</li>\n</ol>\n<p>Also using a very small learning rate. 2e-6.</p>",
      "rawMarkdown": "Some things that helped me.\n\n1. remove weights from loss_fn during training\n2. remove dropout\n3. remove layer norm\n4. remove gradient clipping\n\nAlso using a very small learning rate. 2e-6.",
      "votes": null
    },
    {
      "id": "3069518",
      "postDate": "12/11/2024 15:20:22",
      "content": "<p>Here are a few insights that I've developed in my own training of an NN based model (MLP specific):</p>\n<ol>\n<li>Model Architecture: I've found that wider models outperform deeper models. Max 2 hidden layers and experiment with width. 128 nodes per layer has worked well for me. </li>\n<li>Training Loop: I saw a large improvement in my model when shuffling batches every epoch and cutting off epochs after ~10k batches. Additionally, shuffling your validation batches helps remove bias of the current status of the market on the day you are predicting. </li>\n<li>Small learning rate (~1e-5), large batch sizes (~1024): Since the inputs have quite a bit of noise, large batch sizes help with smoothing especially if you are using any batch normalization. </li>\n<li>Loss function: I've tinkered with the loss function quite a bit. I found that an r2 loss can lead to the local minimum of predicting near zero for every instance (e.g. the trivial minimum) due in part to the sum of true labels being in the denominator. If you truly want to maximize the LB metric, your model should be generating predictions which minimize the term in the numerator of the r2 score. MSE, RMSE, and MAE have all worked well for me. I have not played around with including weights in the loss, but my intuition tells me that this would cause issues based on the amount of noise in the data. </li>\n</ol>",
      "rawMarkdown": "Here are a few insights that I've developed in my own training of an NN based model (MLP specific):\n\n1. Model Architecture: I've found that wider models outperform deeper models. Max 2 hidden layers and experiment with width. 128 nodes per layer has worked well for me. \n2. Training Loop: I saw a large improvement in my model when shuffling batches every epoch and cutting off epochs after ~10k batches. Additionally, shuffling your validation batches helps remove bias of the current status of the market on the day you are predicting. \n3. Small learning rate (~1e-5), large batch sizes (~1024): Since the inputs have quite a bit of noise, large batch sizes help with smoothing especially if you are using any batch normalization. \n4. Loss function: I've tinkered with the loss function quite a bit. I found that an r2 loss can lead to the local minimum of predicting near zero for every instance (e.g. the trivial minimum) due in part to the sum of true labels being in the denominator. If you truly want to maximize the LB metric, your model should be generating predictions which minimize the term in the numerator of the r2 score. MSE, RMSE, and MAE have all worked well for me. I have not played around with including weights in the loss, but my intuition tells me that this would cause issues based on the amount of noise in the data.",
      "votes": null
    },
    {
      "id": "3072532",
      "postDate": "12/15/2024 11:19:47",
      "content": "<p>Hey do you know why shuffle does a better job than normal time split? Time series avoids data leak but gives me a much worse result than normal shuffle.</p>",
      "rawMarkdown": "Hey do you know why shuffle does a better job than normal time split? Time series avoids data leak but gives me a much worse result than normal shuffle.",
      "votes": null
    },
    {
      "id": "3073249",
      "postDate": "12/16/2024 08:22:23",
      "content": "<ol>\n<li>Normalization layers are a trap, you can get good results without them using careful initialization and data prep.</li>\n<li>Activation functions don't matter that much.</li>\n</ol>",
      "rawMarkdown": "1. Normalization layers are a trap, you can get good results without them using careful initialization and data prep.\n2. Activation functions don't matter that much.",
      "votes": null
    },
    {
      "id": "3073896",
      "postDate": "12/17/2024 04:49:15",
      "content": "<p>Removing dropout is surprising to me. Would you just train a smaller network for fewer epochs to avoid overfitting?</p>",
      "rawMarkdown": "Removing dropout is surprising to me. Would you just train a smaller network for fewer epochs to avoid overfitting?",
      "votes": null
    },
    {
      "id": "3073900",
      "postDate": "12/17/2024 04:57:00",
      "content": "<p>its possible. I get wild variation in performance between configs. and im doing a horrible job of remembering what yields what results. Currently trying to replicate performance I had 5 days ago because I forgot what configuration I had </p>",
      "rawMarkdown": "its possible. I get wild variation in performance between configs. and im doing a horrible job of remembering what yields what results. Currently trying to replicate performance I had 5 days ago because I forgot what configuration I had",
      "votes": null
    },
    {
      "id": "3073993",
      "postDate": "12/17/2024 07:37:22",
      "content": "<p>Try using wandb (or a similar tool) to track your experiments. It not only saves your metrics and configs, but also your code (if you’re using colab), git commit hash, and any other information you want to preserve. It has saved me multiple times from losing good models I trained days/weeks ago :)</p>",
      "rawMarkdown": "Try using wandb (or a similar tool) to track your experiments. It not only saves your metrics and configs, but also your code (if you’re using colab), git commit hash, and any other information you want to preserve. It has saved me multiple times from losing good models I trained days/weeks ago :)",
      "votes": null
    },
    {
      "id": "3074012",
      "postDate": "12/17/2024 08:10:22",
      "content": "<p>I've fixed it now. I think I was just severely overfitting. Now I get ~006 with basic fully connected neuralnet, which seems about right, from what I've gathered.</p>",
      "rawMarkdown": "I've fixed it now. I think I was just severely overfitting. Now I get ~006 with basic fully connected neuralnet, which seems about right, from what I've gathered.",
      "votes": null
    },
    {
      "id": "3082144",
      "postDate": "12/27/2024 17:52:05",
      "content": "<p>Hi! Jon! May I ask why the Activation functions does not matter that much? It seems important :)</p>",
      "rawMarkdown": "Hi! Jon! May I ask why the Activation functions does not matter that much? It seems important :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3066248,
      "author_name": "rocstone",
      "author_url": "",
      "post_date": "12/07/2024 21:42:05",
      "content": "<blockquote>\n  <p>I'm just going through the dataset chronologically one time</p>\n</blockquote>\n<p>No, you should shuffle it.</p>\n<p>small lr, simple articheture (do not go deep), weighted r2 score as loss, you should got improvement coz my MLP got 0.0050 score on online test. Let me know hen you got any improvement of result.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3072532,
          "author_name": "yimin218",
          "author_url": "",
          "post_date": "12/15/2024 11:19:47",
          "content": "<p>Hey do you know why shuffle does a better job than normal time split? Time series avoids data leak but gives me a much worse result than normal shuffle.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3068328,
      "author_name": "maciejzawadzki",
      "author_url": "",
      "post_date": "12/10/2024 06:43:16",
      "content": "<p>Something seems to be off, you should be able to get a decent score with a NN model.  What kind of score do you get on a validation portion of the training data -- say the last 120 days?  You'd probably have to share a bit more about what you're doing in order for anyone to give you specific pointers.  There is a public notebook that includes a basic fully connected NN, you should probably check that one out and see whether it produces different results from what you're getting.  I don't think you need anything too fancy on the NN model to get performance in the lasso range IMHO.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3074012,
          "author_name": "reeeeeeeeeeeeeee",
          "author_url": "",
          "post_date": "12/17/2024 08:10:22",
          "content": "<p>I've fixed it now. I think I was just severely overfitting. Now I get ~006 with basic fully connected neuralnet, which seems about right, from what I've gathered.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3069422,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "12/11/2024 13:26:19",
      "content": "<p>Some things that helped me.</p>\n<ol>\n<li>remove weights from loss_fn during training</li>\n<li>remove dropout</li>\n<li>remove layer norm</li>\n<li>remove gradient clipping</li>\n</ol>\n<p>Also using a very small learning rate. 2e-6.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3073896,
          "author_name": "bradywynn",
          "author_url": "",
          "post_date": "12/17/2024 04:49:15",
          "content": "<p>Removing dropout is surprising to me. Would you just train a smaller network for fewer epochs to avoid overfitting?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3073900,
              "author_name": "michaeltimbs",
              "author_url": "",
              "post_date": "12/17/2024 04:57:00",
              "content": "<p>its possible. I get wild variation in performance between configs. and im doing a horrible job of remembering what yields what results. Currently trying to replicate performance I had 5 days ago because I forgot what configuration I had </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3073993,
                  "author_name": "eivolkova",
                  "author_url": "",
                  "post_date": "12/17/2024 07:37:22",
                  "content": "<p>Try using wandb (or a similar tool) to track your experiments. It not only saves your metrics and configs, but also your code (if you’re using colab), git commit hash, and any other information you want to preserve. It has saved me multiple times from losing good models I trained days/weeks ago :)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3069518,
      "author_name": "bryancrossman",
      "author_url": "",
      "post_date": "12/11/2024 15:20:22",
      "content": "<p>Here are a few insights that I've developed in my own training of an NN based model (MLP specific):</p>\n<ol>\n<li>Model Architecture: I've found that wider models outperform deeper models. Max 2 hidden layers and experiment with width. 128 nodes per layer has worked well for me. </li>\n<li>Training Loop: I saw a large improvement in my model when shuffling batches every epoch and cutting off epochs after ~10k batches. Additionally, shuffling your validation batches helps remove bias of the current status of the market on the day you are predicting. </li>\n<li>Small learning rate (~1e-5), large batch sizes (~1024): Since the inputs have quite a bit of noise, large batch sizes help with smoothing especially if you are using any batch normalization. </li>\n<li>Loss function: I've tinkered with the loss function quite a bit. I found that an r2 loss can lead to the local minimum of predicting near zero for every instance (e.g. the trivial minimum) due in part to the sum of true labels being in the denominator. If you truly want to maximize the LB metric, your model should be generating predictions which minimize the term in the numerator of the r2 score. MSE, RMSE, and MAE have all worked well for me. I have not played around with including weights in the loss, but my intuition tells me that this would cause issues based on the amount of noise in the data. </li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3073249,
      "author_name": "probablynobody",
      "author_url": "",
      "post_date": "12/16/2024 08:22:23",
      "content": "<ol>\n<li>Normalization layers are a trap, you can get good results without them using careful initialization and data prep.</li>\n<li>Activation functions don't matter that much.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 3082144,
          "author_name": "larrylin666",
          "author_url": "",
          "post_date": "12/27/2024 17:52:05",
          "content": "<p>Hi! Jon! May I ask why the Activation functions does not matter that much? It seems important :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3066038": "So far, all my attempts at using neural networks have yielded far worse performance than any GBDT based model, or even linear regression. It ends up generalizing catastrophically to my validation.\n\nIf I use something like lasso regression, I get a validation score of like 0.0035. If I use neural networks I get validation scores -10.\n\nI use a basic NN architecture. Dense, feedforward, GELU, dropout, layernorm, tuned with Adam. I've tried to do hyperparameter search over basically all the parameters (learning rate, dropout rate, l2 norm, momentum parameters, number of layers, dimension of hidden layers).\n\nI'm wondering if I'm doing too few gradient steps per batch? I'm just going through the dataset chronologically one time. Should I be doing more than one gradient step per sample in the dataset?",
    "3066248": "> I'm just going through the dataset chronologically one time\n\nNo, you should shuffle it.\n\nsmall lr, simple articheture (do not go deep), weighted r2 score as loss, you should got improvement coz my MLP got 0.0050 score on online test. Let me know hen you got any improvement of result.",
    "3068328": "Something seems to be off, you should be able to get a decent score with a NN model.  What kind of score do you get on a validation portion of the training data -- say the last 120 days?  You'd probably have to share a bit more about what you're doing in order for anyone to give you specific pointers.  There is a public notebook that includes a basic fully connected NN, you should probably check that one out and see whether it produces different results from what you're getting.  I don't think you need anything too fancy on the NN model to get performance in the lasso range IMHO.",
    "3069422": "Some things that helped me.\n\n1. remove weights from loss_fn during training\n2. remove dropout\n3. remove layer norm\n4. remove gradient clipping\n\nAlso using a very small learning rate. 2e-6.",
    "3069518": "Here are a few insights that I've developed in my own training of an NN based model (MLP specific):\n\n1. Model Architecture: I've found that wider models outperform deeper models. Max 2 hidden layers and experiment with width. 128 nodes per layer has worked well for me. \n2. Training Loop: I saw a large improvement in my model when shuffling batches every epoch and cutting off epochs after ~10k batches. Additionally, shuffling your validation batches helps remove bias of the current status of the market on the day you are predicting. \n3. Small learning rate (~1e-5), large batch sizes (~1024): Since the inputs have quite a bit of noise, large batch sizes help with smoothing especially if you are using any batch normalization. \n4. Loss function: I've tinkered with the loss function quite a bit. I found that an r2 loss can lead to the local minimum of predicting near zero for every instance (e.g. the trivial minimum) due in part to the sum of true labels being in the denominator. If you truly want to maximize the LB metric, your model should be generating predictions which minimize the term in the numerator of the r2 score. MSE, RMSE, and MAE have all worked well for me. I have not played around with including weights in the loss, but my intuition tells me that this would cause issues based on the amount of noise in the data.",
    "3072532": "Hey do you know why shuffle does a better job than normal time split? Time series avoids data leak but gives me a much worse result than normal shuffle.",
    "3073249": "1. Normalization layers are a trap, you can get good results without them using careful initialization and data prep.\n2. Activation functions don't matter that much.",
    "3073896": "Removing dropout is surprising to me. Would you just train a smaller network for fewer epochs to avoid overfitting?",
    "3073900": "its possible. I get wild variation in performance between configs. and im doing a horrible job of remembering what yields what results. Currently trying to replicate performance I had 5 days ago because I forgot what configuration I had",
    "3073993": "Try using wandb (or a similar tool) to track your experiments. It not only saves your metrics and configs, but also your code (if you’re using colab), git commit hash, and any other information you want to preserve. It has saved me multiple times from losing good models I trained days/weeks ago :)",
    "3074012": "I've fixed it now. I think I was just severely overfitting. Now I get ~006 with basic fully connected neuralnet, which seems about right, from what I've gathered.",
    "3082144": "Hi! Jon! May I ask why the Activation functions does not matter that much? It seems important :)"
  },
  "source": "meta"
}