{
  "id": 582305,
  "title": "RNN in crypto market prediction",
  "url": "/competitions/drw-crypto-market-prediction/discussion/582305",
  "author_name": "Dmitry Kiryukhin",
  "post_date": "2025-05-30T10:28:57.405000",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I'm trying to find the RNN architecture for making a forecast, but the maximum I set using RNN was about 0.06 Pearson. In your opinion, which architecture is better suited? I started by using LSTM and transmitting data to it in a time series format with a batch_size of 1440 (equal to the size of a day in the sample), using a few linear layers and LeakyReLU as activation functions to avoid loss and attenuation of the gradient. Then I started trying the usual linear layer with the same batch_size, but the result still didn't meet expectations. Next, I complicated the model and added skip connections. The result did not exceed the LSTM with three linear layers. So which architecture is most preferred? And how did you draw the model's attention to the small seasonality of the data (such as trading day, holidays, etc.)</p>",
  "messages": [
    {
      "id": 3213702,
      "postDate": "2025-05-30T10:51:32.537Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/poitrew\" target=\"_blank\">@poitrew</a>, instead of jumping straight to such computationally expensive techniques, I'd recommend trying some conventional machine learning methods first.</p>",
      "rawMarkdown": "Hey @poitrew, instead of jumping straight to such computationally expensive techniques, I'd recommend trying some conventional machine learning methods first.",
      "votes": 3
    },
    {
      "id": 3213720,
      "postDate": "2025-05-30T11:07:45.650Z",
      "content": "<p>Hi, how could you apply time series NN on this competition since the timestamp of test data is masked?</p>\n<p>BTW, the standardization for data does not work better for training from my test.</p>",
      "rawMarkdown": "Hi, how could you apply time series NN on this competition since the timestamp of test data is masked?\n\nBTW, the standardization for data does not work better for training from my test.",
      "votes": 2,
      "replies": [
        {
          "id": 3214919,
          "postDate": "2025-06-01T09:14:29.223Z",
          "content": "<p>I delete normalization, add some epochs, some dense layers and pearson coef on test set increase. Batch_size was 256. Thank for your note</p>",
          "rawMarkdown": "I delete normalization, add some epochs, some dense layers and pearson coef on test set increase. Batch_size was 256. Thank for your note"
        }
      ]
    },
    {
      "id": 3213687,
      "postDate": "2025-05-30T10:28:57.407Z",
      "content": "<p>I'm trying to find the RNN architecture for making a forecast, but the maximum I set using RNN was about 0.06 Pearson. In your opinion, which architecture is better suited? I started by using LSTM and transmitting data to it in a time series format with a batch_size of 1440 (equal to the size of a day in the sample), using a few linear layers and LeakyReLU as activation functions to avoid loss and attenuation of the gradient. Then I started trying the usual linear layer with the same batch_size, but the result still didn't meet expectations. Next, I complicated the model and added skip connections. The result did not exceed the LSTM with three linear layers. So which architecture is most preferred? And how did you draw the model's attention to the small seasonality of the data (such as trading day, holidays, etc.)</p>",
      "rawMarkdown": "I'm trying to find the RNN architecture for making a forecast, but the maximum I set using RNN was about 0.06 Pearson. In your opinion, which architecture is better suited? I started by using LSTM and transmitting data to it in a time series format with a batch_size of 1440 (equal to the size of a day in the sample), using a few linear layers and LeakyReLU as activation functions to avoid loss and attenuation of the gradient. Then I started trying the usual linear layer with the same batch_size, but the result still didn't meet expectations. Next, I complicated the model and added skip connections. The result did not exceed the LSTM with three linear layers. So which architecture is most preferred? And how did you draw the model's attention to the small seasonality of the data (such as trading day, holidays, etc.)",
      "votes": 2
    },
    {
      "id": 3215114,
      "postDate": "2025-06-01T15:56:03.880Z",
      "content": "<p>Before doing that, I highly suggest trying more conventional models. The timestamp of test data is masked so it will be challenging to implement a time series NN such as RNN, or geting any time features (e.g., hour). </p>",
      "rawMarkdown": "Before doing that, I highly suggest trying more conventional models. The timestamp of test data is masked so it will be challenging to implement a time series NN such as RNN, or geting any time features (e.g., hour). "
    },
    {
      "id": 3213689,
      "postDate": "2025-05-30T10:32:15.917Z",
      "content": "<p>It should be noted that I converted the data to the range [0;1], after which I used the division into bath without shuffle</p>\n<p>def residual_block(x, units, dropout_rate, alpha=0.01):<br>\n    y = Dense(units)(x)<br>\n    y = BatchNormalization()(y)<br>\n    y = LeakyReLU(alpha=alpha)(y)<br>\n    y = Dropout(dropout_rate)(y)<br>\n    if x.shape[-1] != units:<br>\n        x = Dense(units)(x)<br>\n        x = BatchNormalization()(x)</p>\n<pre><code> = Add()([x, y])\n = LeakyReLU(=)()\n \n</code></pre>\n<p>inputs = Input(shape=(840,))<br>\nx = Dense(512)(inputs)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.2)(x)</p>\n<p>x = residual_block(x, 512, 0.15)<br>\nx = residual_block(x, 512, 0.15)</p>\n<p>x = Dense(256)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.2)(x)</p>\n<p>x = residual_block(x, 256, 0.15)<br>\nx = residual_block(x, 256, 0.15)</p>\n<p>x = Dense(128)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.15)(x)</p>\n<p>x = residual_block(x, 128, 0.1)</p>\n<p>x = Dense(96)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.1)(x)</p>\n<p>x = residual_block(x, 96, 0.1)</p>\n<p>x = Dense(64)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.1)(x)</p>\n<p>x = Dense(32)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.05)(x)<br>\noutputs = Dense(1, dtype='float32')(x)<br>\nMy best LB score model</p>",
      "rawMarkdown": "It should be noted that I converted the data to the range [0;1], after which I used the division into bath without shuffle\n\ndef residual_block(x, units, dropout_rate, alpha=0.01):\n    y = Dense(units)(x)\n    y = BatchNormalization()(y)\n    y = LeakyReLU(alpha=alpha)(y)\n    y = Dropout(dropout_rate)(y)\n    if x.shape[-1] != units:\n        x = Dense(units)(x)\n        x = BatchNormalization()(x)\n    \n    out = Add()([x, y])\n    out = LeakyReLU(alpha=alpha)(out)\n    return out\n\ninputs = Input(shape=(840,))\nx = Dense(512)(inputs)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.2)(x)\n\nx = residual_block(x, 512, 0.15)\nx = residual_block(x, 512, 0.15)\n\nx = Dense(256)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.2)(x)\n\nx = residual_block(x, 256, 0.15)\nx = residual_block(x, 256, 0.15)\n\nx = Dense(128)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.15)(x)\n\nx = residual_block(x, 128, 0.1)\n\nx = Dense(96)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.1)(x)\n\nx = residual_block(x, 96, 0.1)\n\nx = Dense(64)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.1)(x)\n\nx = Dense(32)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.05)(x)\noutputs = Dense(1, dtype='float32')(x)\nMy best LB score model",
      "replies": [
        {
          "id": 3213792,
          "postDate": "2025-05-30T13:13:18.830Z",
          "content": "<p>Are you using MSE or Pearson Correlation for loss function? </p>",
          "rawMarkdown": "Are you using MSE or Pearson Correlation for loss function? ",
          "votes": 1,
          "replies": [
            {
              "id": 3214917,
              "postDate": "2025-06-01T09:13:06.740Z",
              "content": "<p>I used MSE. Add pearson correlation like metrics function. But I delet MinMaxScaler + add some dense for my NN and Pearson coef on test set increase</p>",
              "rawMarkdown": "I used MSE. Add pearson correlation like metrics function. But I delet MinMaxScaler + add some dense for my NN and Pearson coef on test set increase"
            }
          ]
        }
      ]
    },
    {
      "id": 3214047,
      "postDate": "2025-05-30T20:52:42.987Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3213702,
      "author_name": "Vishal Painjane",
      "author_url": "",
      "post_date": "2025-05-30T10:51:32.537000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/poitrew\" target=\"_blank\">@poitrew</a>, instead of jumping straight to such computationally expensive techniques, I'd recommend trying some conventional machine learning methods first.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3213720,
      "author_name": "Yourui Wang",
      "author_url": "",
      "post_date": "2025-05-30T11:07:45.650000",
      "content": "<p>Hi, how could you apply time series NN on this competition since the timestamp of test data is masked?</p>\n<p>BTW, the standardization for data does not work better for training from my test.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3214919,
          "author_name": "Dmitry Kiryukhin",
          "author_url": "",
          "post_date": "2025-06-01T09:14:29.223000",
          "content": "<p>I delete normalization, add some epochs, some dense layers and pearson coef on test set increase. Batch_size was 256. Thank for your note</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3215114,
      "author_name": "SergioMiguelM",
      "author_url": "",
      "post_date": "2025-06-01T15:56:03.880000",
      "content": "<p>Before doing that, I highly suggest trying more conventional models. The timestamp of test data is masked so it will be challenging to implement a time series NN such as RNN, or geting any time features (e.g., hour). </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3213689,
      "author_name": "Dmitry Kiryukhin",
      "author_url": "",
      "post_date": "2025-05-30T10:32:15.917000",
      "content": "<p>It should be noted that I converted the data to the range [0;1], after which I used the division into bath without shuffle</p>\n<p>def residual_block(x, units, dropout_rate, alpha=0.01):<br>\n    y = Dense(units)(x)<br>\n    y = BatchNormalization()(y)<br>\n    y = LeakyReLU(alpha=alpha)(y)<br>\n    y = Dropout(dropout_rate)(y)<br>\n    if x.shape[-1] != units:<br>\n        x = Dense(units)(x)<br>\n        x = BatchNormalization()(x)</p>\n<pre><code> = Add()([x, y])\n = LeakyReLU(=)()\n \n</code></pre>\n<p>inputs = Input(shape=(840,))<br>\nx = Dense(512)(inputs)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.2)(x)</p>\n<p>x = residual_block(x, 512, 0.15)<br>\nx = residual_block(x, 512, 0.15)</p>\n<p>x = Dense(256)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.2)(x)</p>\n<p>x = residual_block(x, 256, 0.15)<br>\nx = residual_block(x, 256, 0.15)</p>\n<p>x = Dense(128)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.15)(x)</p>\n<p>x = residual_block(x, 128, 0.1)</p>\n<p>x = Dense(96)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.1)(x)</p>\n<p>x = residual_block(x, 96, 0.1)</p>\n<p>x = Dense(64)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.1)(x)</p>\n<p>x = Dense(32)(x)<br>\nx = BatchNormalization()(x)<br>\nx = LeakyReLU(alpha=0.01)(x)<br>\nx = Dropout(0.05)(x)<br>\noutputs = Dense(1, dtype='float32')(x)<br>\nMy best LB score model</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3213792,
          "author_name": "byunjins",
          "author_url": "",
          "post_date": "2025-05-30T13:13:18.830000",
          "content": "<p>Are you using MSE or Pearson Correlation for loss function? </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3214917,
              "author_name": "Dmitry Kiryukhin",
              "author_url": "",
              "post_date": "2025-06-01T09:13:06.740000",
              "content": "<p>I used MSE. Add pearson correlation like metrics function. But I delet MinMaxScaler + add some dense for my NN and Pearson coef on test set increase</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3214047,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-30T20:52:42.987000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3213702": "Hey @poitrew, instead of jumping straight to such computationally expensive techniques, I'd recommend trying some conventional machine learning methods first.",
    "3213720": "Hi, how could you apply time series NN on this competition since the timestamp of test data is masked?\n\nBTW, the standardization for data does not work better for training from my test.",
    "3213687": "I'm trying to find the RNN architecture for making a forecast, but the maximum I set using RNN was about 0.06 Pearson. In your opinion, which architecture is better suited? I started by using LSTM and transmitting data to it in a time series format with a batch_size of 1440 (equal to the size of a day in the sample), using a few linear layers and LeakyReLU as activation functions to avoid loss and attenuation of the gradient. Then I started trying the usual linear layer with the same batch_size, but the result still didn't meet expectations. Next, I complicated the model and added skip connections. The result did not exceed the LSTM with three linear layers. So which architecture is most preferred? And how did you draw the model's attention to the small seasonality of the data (such as trading day, holidays, etc.)",
    "3215114": "Before doing that, I highly suggest trying more conventional models. The timestamp of test data is masked so it will be challenging to implement a time series NN such as RNN, or geting any time features (e.g., hour). ",
    "3213689": "It should be noted that I converted the data to the range [0;1], after which I used the division into bath without shuffle\n\ndef residual_block(x, units, dropout_rate, alpha=0.01):\n    y = Dense(units)(x)\n    y = BatchNormalization()(y)\n    y = LeakyReLU(alpha=alpha)(y)\n    y = Dropout(dropout_rate)(y)\n    if x.shape[-1] != units:\n        x = Dense(units)(x)\n        x = BatchNormalization()(x)\n    \n    out = Add()([x, y])\n    out = LeakyReLU(alpha=alpha)(out)\n    return out\n\ninputs = Input(shape=(840,))\nx = Dense(512)(inputs)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.2)(x)\n\nx = residual_block(x, 512, 0.15)\nx = residual_block(x, 512, 0.15)\n\nx = Dense(256)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.2)(x)\n\nx = residual_block(x, 256, 0.15)\nx = residual_block(x, 256, 0.15)\n\nx = Dense(128)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.15)(x)\n\nx = residual_block(x, 128, 0.1)\n\nx = Dense(96)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.1)(x)\n\nx = residual_block(x, 96, 0.1)\n\nx = Dense(64)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.1)(x)\n\nx = Dense(32)(x)\nx = BatchNormalization()(x)\nx = LeakyReLU(alpha=0.01)(x)\nx = Dropout(0.05)(x)\noutputs = Dense(1, dtype='float32')(x)\nMy best LB score model",
    "3214047": ""
  }
}