{
  "id": 547644,
  "title": "Discussion prompts for the week of 25.11 - 01.12",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/547644",
  "author_name": "Victor Shlepov",
  "post_date": "2024-11-22T20:59:26.965000",
  "votes": 7,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Hey folks, it was a good week, right? The top 10 is getting denser now, which is actually a positive sign - the more competition, the better! Let’s choose a topic for next week to discuss. No secret recipes - keep those to yourself (I don’t believe in any for the financial markets anyway) - just some high-level observations, insights, and ideas. Let’s make this competition a little more interactive and live!<br>\nSo, what’s your call? Normalization, online learning, lags, augmentation strategy, or something else? Just drop your suggestions here. We’ll pick one, and I’ll share my thoughts; I hope you’ll do the same too!</p>",
  "messages": [
    {
      "id": 3052807,
      "postDate": "2024-11-22T20:59:26.967Z",
      "content": "<p>Hey folks, it was a good week, right? The top 10 is getting denser now, which is actually a positive sign - the more competition, the better! Let’s choose a topic for next week to discuss. No secret recipes - keep those to yourself (I don’t believe in any for the financial markets anyway) - just some high-level observations, insights, and ideas. Let’s make this competition a little more interactive and live!<br>\nSo, what’s your call? Normalization, online learning, lags, augmentation strategy, or something else? Just drop your suggestions here. We’ll pick one, and I’ll share my thoughts; I hope you’ll do the same too!</p>",
      "rawMarkdown": "Hey folks, it was a good week, right? The top 10 is getting denser now, which is actually a positive sign - the more competition, the better! Let’s choose a topic for next week to discuss. No secret recipes - keep those to yourself (I don’t believe in any for the financial markets anyway) - just some high-level observations, insights, and ideas. Let’s make this competition a little more interactive and live!\nSo, what’s your call? Normalization, online learning, lags, augmentation strategy, or something else? Just drop your suggestions here. We’ll pick one, and I’ll share my thoughts; I hope you’ll do the same too!",
      "votes": 6
    },
    {
      "id": 3052880,
      "postDate": "2024-11-22T23:50:00.530Z",
      "content": "<p>Normalization. online learning in that order.</p>",
      "rawMarkdown": "Normalization. online learning in that order.",
      "votes": 3
    },
    {
      "id": 3055706,
      "postDate": "2024-11-26T03:29:17.857Z",
      "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> can u share a simple baseline of online learning (NN) ? no matter how simple it is, I tried to build one, but got tons of bug😭</p>",
      "rawMarkdown": "@victorshlepov can u share a simple baseline of online learning (NN) ? no matter how simple it is, I tried to build one, but got tons of bug😭",
      "votes": 1
    },
    {
      "id": 3054940,
      "postDate": "2024-11-25T10:54:39.790Z",
      "content": "<p>Non-stationary property is the key to this competition, as I noticed. The most important issue is, how to <strong>fit a general model</strong> and to <strong>adjust it flexibly</strong> according to the different pattern. The big problem is, we are limited to the time constraint and submission form in batch instead of time series of api. How to track this problem? 1. Form the data as data streaming with feature engineering or 2. form the data as time series with LSTM or Transformer to capture the non-stationary. Do you have any advice?</p>",
      "rawMarkdown": "Non-stationary property is the key to this competition, as I noticed. The most important issue is, how to **fit a general model** and to **adjust it flexibly** according to the different pattern. The big problem is, we are limited to the time constraint and submission form in batch instead of time series of api. How to track this problem? 1. Form the data as data streaming with feature engineering or 2. form the data as time series with LSTM or Transformer to capture the non-stationary. Do you have any advice?",
      "votes": 2,
      "replies": [
        {
          "id": 3054941,
          "postDate": "2024-11-25T10:56:58.507Z",
          "content": "<p>An interesting observation is that model trained by large dataset probably perform worse than the model with small dataset but in a stationary pattern.</p>",
          "rawMarkdown": "An interesting observation is that model trained by large dataset probably perform worse than the model with small dataset but in a stationary pattern.",
          "votes": 1
        },
        {
          "id": 3054943,
          "postDate": "2024-11-25T11:01:48.600Z",
          "content": "<p>what about bagging by combining multiple models (Purely time series model based on lags + a linear model such as LASSO + a non Linear model), WDYT? The challenge would be how to distribute the weights among the predictions produced by the different models</p>",
          "rawMarkdown": "what about bagging by combining multiple models (Purely time series model based on lags + a linear model such as LASSO + a non Linear model), WDYT? The challenge would be how to distribute the weights among the predictions produced by the different models",
          "replies": [
            {
              "id": 3054954,
              "postDate": "2024-11-25T11:17:54.503Z",
              "content": "<p>We have to figure out the non-stationary pattern at first to assign different model to fit it and assess the similarity between test data and one pattern to make synthetic weight.</p>",
              "rawMarkdown": "We have to figure out the non-stationary pattern at first to assign different model to fit it and assess the similarity between test data and one pattern to make synthetic weight."
            },
            {
              "id": 3054958,
              "postDate": "2024-11-25T11:23:57.230Z",
              "content": "<p>The first solution is use statistical model to analyze it, as paper \"Multi-dimensional latent group structures with heterogeneous distributions\" here: <a href=\"https://www.sciencedirect.com/science/article/pii/S0304407621002177\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S0304407621002177</a>. Or we can just use Transformer to capture it automatically as paper \"Transformers as Statisticians: Provable In-Context<br>\nLearning with In-Context Algorithm Selection\" said here: <a href=\"https://proceedings.neurips.cc/paper_files/paper/2023/file/b2e63e36c57e153b9015fece2352a9f9-Paper-Conference.pdf\" target=\"_blank\">https://proceedings.neurips.cc/paper_files/paper/2023/file/b2e63e36c57e153b9015fece2352a9f9-Paper-Conference.pdf</a> 😂</p>",
              "rawMarkdown": "The first solution is use statistical model to analyze it, as paper \"Multi-dimensional latent group structures with heterogeneous distributions\" here: https://www.sciencedirect.com/science/article/pii/S0304407621002177. Or we can just use Transformer to capture it automatically as paper \"Transformers as Statisticians: Provable In-Context\nLearning with In-Context Algorithm Selection\" said here: https://proceedings.neurips.cc/paper_files/paper/2023/file/b2e63e36c57e153b9015fece2352a9f9-Paper-Conference.pdf 😂"
            }
          ]
        },
        {
          "id": 3086912,
          "postDate": "2025-01-02T20:39:06.310Z",
          "content": "<p>All features and target are stationary</p>",
          "rawMarkdown": "All features and target are stationary"
        }
      ]
    },
    {
      "id": 3054590,
      "postDate": "2024-11-24T21:10:25.540Z",
      "content": "<p>Oh, but I had a quick question?</p>",
      "rawMarkdown": "Oh, but I had a quick question?",
      "replies": [
        {
          "id": 3054646,
          "postDate": "2024-11-25T00:02:06.530Z",
          "content": "<p>I think it's better ask it then wasting a chance to get the answer :)</p>",
          "rawMarkdown": "I think it's better ask it then wasting a chance to get the answer :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 3086743,
      "postDate": "2025-01-02T16:32:11.910Z",
      "content": "<p>Could you please provide insights or strategies for reducing the discrepancy between my local CV scores and the leaderboard (LB) scores? My main challenge seems to be that my local CV setup does not reliably indicate which models will perform better on the LB. Any advice on improving CV strategies or diagnosing this issue would be greatly appreciated.</p>",
      "rawMarkdown": "Could you please provide insights or strategies for reducing the discrepancy between my local CV scores and the leaderboard (LB) scores? My main challenge seems to be that my local CV setup does not reliably indicate which models will perform better on the LB. Any advice on improving CV strategies or diagnosing this issue would be greatly appreciated."
    },
    {
      "id": 3086718,
      "postDate": "2025-01-02T16:10:16.347Z",
      "content": "<p>new feature idea:  considering that the total sum of weights is stable each day across each time_id, do you think it's worth calculating for each time_id,  the weight as a % of weights sum ? </p>",
      "rawMarkdown": "new feature idea:  considering that the total sum of weights is stable each day across each time_id, do you think it's worth calculating for each time_id,  the weight as a % of weights sum ? \n\n\n"
    },
    {
      "id": 3056111,
      "postDate": "2024-11-26T15:11:31.977Z",
      "content": "<p>I finally got my first ensemble trained and submitted only to drop from 0.0034 baseline to -0.006 with the ensemble so that was fun!</p>",
      "rawMarkdown": "I finally got my first ensemble trained and submitted only to drop from 0.0034 baseline to -0.006 with the ensemble so that was fun!"
    },
    {
      "id": 3053871,
      "postDate": "2024-11-24T04:03:58.063Z",
      "content": "<p>Can you talk about the use of tags in this competition?</p>",
      "rawMarkdown": "Can you talk about the use of tags in this competition?",
      "replies": [
        {
          "id": 3054304,
          "postDate": "2024-11-24T14:47:31.350Z",
          "content": "<p>I tried to analyze the tags closely and I noticed that these are used to group the features (same for the responders tags) into features (or responders) that have the same behavior/characteristics. <br>\nAlso from what I have understood, the tags for the features have nothing to do with the tags for the responders.</p>",
          "rawMarkdown": "I tried to analyze the tags closely and I noticed that these are used to group the features (same for the responders tags) into features (or responders) that have the same behavior/characteristics. \nAlso from what I have understood, the tags for the features have nothing to do with the tags for the responders.",
          "votes": 2,
          "replies": [
            {
              "id": 3054509,
              "postDate": "2024-11-24T18:42:20.010Z",
              "content": "<p>Agree. On top that - tag is some sort of a constant, right? And since tags are the same for every single time step - I'd presume that neural network is capable of learning this kind of representation on it's own (if it's a useful one). So, I don't think it worth spending time and effort solving the \"tags quiz\", better to focus on a meaningful \"moving\" components instead… </p>",
              "rawMarkdown": "Agree. On top that - tag is some sort of a constant, right? And since tags are the same for every single time step - I'd presume that neural network is capable of learning this kind of representation on it's own (if it's a useful one). So, I don't think it worth spending time and effort solving the \"tags quiz\", better to focus on a meaningful \"moving\" components instead... ",
              "votes": 3
            },
            {
              "id": 3054533,
              "postDate": "2024-11-24T19:27:33.347Z",
              "content": "<p>Fully agree with you.</p>",
              "rawMarkdown": "Fully agree with you."
            }
          ]
        }
      ]
    },
    {
      "id": 3053809,
      "postDate": "2024-11-24T00:03:11.847Z",
      "content": "<p>Have you done any feature engineering?<br>\nand by feature engineering I mean adding new features from the existing one </p>",
      "rawMarkdown": "Have you done any feature engineering?\nand by feature engineering I mean adding new features from the existing one ",
      "replies": [
        {
          "id": 3053811,
          "postDate": "2024-11-24T00:09:29.147Z",
          "content": "<p>No, I think this is a dead end. The neural network would likely outperform any of us in \"feature engineering.\"<br>\nAnd in this particular domain feature engineering is twice as useless. It's kind of \"silver bullet\" story - to come up with an artificial feature that would serve as a predictor for the market…</p>",
          "rawMarkdown": "No, I think this is a dead end. The neural network would likely outperform any of us in \"feature engineering.\"\nAnd in this particular domain feature engineering is twice as useless. It's kind of \"silver bullet\" story - to come up with an artificial feature that would serve as a predictor for the market...",
          "votes": 1,
          "replies": [
            {
              "id": 3053813,
              "postDate": "2024-11-24T00:18:38.920Z",
              "content": "<p>In some time series problems feature engineering can add great value but sure no silver bullet you need to add good set of features <br>\nMaybe this competition is different on that regard for some reasons such as:<br>\n1- the nature of finicial market <br>\n2- some features are maybe already engineered features from other features that might or mightn't exist in the dataset <br>\nSo if you decide to do feature engineering you will do it totally blind </p>",
              "rawMarkdown": "In some time series problems feature engineering can add great value but sure no silver bullet you need to add good set of features \nMaybe this competition is different on that regard for some reasons such as:\n1- the nature of finicial market \n2- some features are maybe already engineered features from other features that might or mightn't exist in the dataset \nSo if you decide to do feature engineering you will do it totally blind "
            }
          ]
        }
      ]
    },
    {
      "id": 3053277,
      "postDate": "2024-11-23T11:28:56.140Z",
      "content": "<p>Do you think we should make different preprocessing for different symbol_id data?</p>",
      "rawMarkdown": "Do you think we should make different preprocessing for different symbol_id data?",
      "replies": [
        {
          "id": 3053315,
          "postDate": "2024-11-23T11:52:51.700Z",
          "content": "<p>I don't know, why would you do that? In finance, we have a kind of a uniform scales - returns, prices, etc. You should have a strong reason to put symbols onto different scales, I guess…</p>",
          "rawMarkdown": "I don't know, why would you do that? In finance, we have a kind of a uniform scales - returns, prices, etc. You should have a strong reason to put symbols onto different scales, I guess...",
          "votes": 2
        },
        {
          "id": 3054926,
          "postDate": "2024-11-25T10:44:01.643Z",
          "content": "<p>To utilize symbol_id's heterogeneity, the submission will be timeout😂 and it's possible to add new features in test data.</p>",
          "rawMarkdown": "To utilize symbol_id's heterogeneity, the submission will be timeout😂 and it's possible to add new features in test data."
        }
      ]
    },
    {
      "id": 3053238,
      "postDate": "2024-11-23T10:07:28.343Z",
      "content": "<p>Let's talk about normalization. Is there a good normalization for time series?</p>",
      "rawMarkdown": "Let's talk about normalization. Is there a good normalization for time series?",
      "replies": [
        {
          "id": 3053241,
          "postDate": "2024-11-23T10:16:19.963Z",
          "content": "<p>I would add \"non stationary time series\" on top of that :)</p>",
          "rawMarkdown": "I would add \"non stationary time series\" on top of that :)",
          "votes": 2,
          "replies": [
            {
              "id": 3053391,
              "postDate": "2024-11-23T12:42:55.457Z",
              "content": "<p>Any tips on how to deal with non-stationary time series?<br>\nLSTM? GRU? Scaling?</p>",
              "rawMarkdown": "Any tips on how to deal with non-stationary time series?\nLSTM? GRU? Scaling?"
            },
            {
              "id": 3053796,
              "postDate": "2024-11-23T23:38:09.973Z",
              "content": "<p>Look, without getting into unnecessary details- there's no \"magic bullet\" here. I'd rather put all this fancy acronyms aside and try to think in some real, industry-specific terms - what and why you want the model to mimic. That's exactly what I do, cause tweaking the code or juggling the different layers with no big picture in mind is no fun at all…</p>",
              "rawMarkdown": "Look, without getting into unnecessary details- there's no \"magic bullet\" here. I'd rather put all this fancy acronyms aside and try to think in some real, industry-specific terms - what and why you want the model to mimic. That's exactly what I do, cause tweaking the code or juggling the different layers with no big picture in mind is no fun at all...",
              "votes": 4
            },
            {
              "id": 3054300,
              "postDate": "2024-11-24T14:43:27.553Z",
              "content": "<p>Thx.<br>\nDo you have any materials to better understand this problem and “industry-specific terms”?</p>",
              "rawMarkdown": "Thx.\nDo you have any materials to better understand this problem and “industry-specific terms”?"
            },
            {
              "id": 3059652,
              "postDate": "2024-12-01T00:34:56.370Z",
              "content": "<p>Hi, thank you so much for your insightful response! I'm not from the finance industry, so I’m not very familiar with the specific use cases in this domain. Could you provide an example or recommend some resources to help me better understand how to build models in finance? Specifically, how do you select appropriate features and understand the applications of different models for market prediction?</p>\n<p>Looking forward to your suggestions, and thanks again!</p>",
              "rawMarkdown": "Hi, thank you so much for your insightful response! I'm not from the finance industry, so I’m not very familiar with the specific use cases in this domain. Could you provide an example or recommend some resources to help me better understand how to build models in finance? Specifically, how do you select appropriate features and understand the applications of different models for market prediction?\n\nLooking forward to your suggestions, and thanks again!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3052821,
      "postDate": "2024-11-22T21:17:02.133Z",
      "content": "<p>How about the topic is whichever of those was the driver behind you climbing to 3rd place so fast? 😝</p>",
      "rawMarkdown": "How about the topic is whichever of those was the driver behind you climbing to 3rd place so fast? 😝",
      "replies": [
        {
          "id": 3053087,
          "postDate": "2024-11-23T06:08:33.713Z",
          "content": "<p>I just stick to my own advice and keep circling back to the same ideas over and over again… :)</p>",
          "rawMarkdown": "I just stick to my own advice and keep circling back to the same ideas over and over again... :)",
          "votes": 4
        }
      ]
    },
    {
      "id": 3054748,
      "postDate": "2024-11-25T05:16:36.603Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3054299,
      "postDate": "2024-11-24T14:41:55.587Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3052880,
      "author_name": "Yax Akaba",
      "author_url": "",
      "post_date": "2024-11-22T23:50:00.530000",
      "content": "<p>Normalization. online learning in that order.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3055706,
      "author_name": "Wayne_127",
      "author_url": "",
      "post_date": "2024-11-26T03:29:17.857000",
      "content": "<p><a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> can u share a simple baseline of online learning (NN) ? no matter how simple it is, I tried to build one, but got tons of bug😭</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3054940,
      "author_name": "#1BuBu",
      "author_url": "",
      "post_date": "2024-11-25T10:54:39.790000",
      "content": "<p>Non-stationary property is the key to this competition, as I noticed. The most important issue is, how to <strong>fit a general model</strong> and to <strong>adjust it flexibly</strong> according to the different pattern. The big problem is, we are limited to the time constraint and submission form in batch instead of time series of api. How to track this problem? 1. Form the data as data streaming with feature engineering or 2. form the data as time series with LSTM or Transformer to capture the non-stationary. Do you have any advice?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3054941,
          "author_name": "#1BuBu",
          "author_url": "",
          "post_date": "2024-11-25T10:56:58.507000",
          "content": "<p>An interesting observation is that model trained by large dataset probably perform worse than the model with small dataset but in a stationary pattern.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3054943,
          "author_name": "benchouikha01",
          "author_url": "",
          "post_date": "2024-11-25T11:01:48.600000",
          "content": "<p>what about bagging by combining multiple models (Purely time series model based on lags + a linear model such as LASSO + a non Linear model), WDYT? The challenge would be how to distribute the weights among the predictions produced by the different models</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3054954,
              "author_name": "#1BuBu",
              "author_url": "",
              "post_date": "2024-11-25T11:17:54.503000",
              "content": "<p>We have to figure out the non-stationary pattern at first to assign different model to fit it and assess the similarity between test data and one pattern to make synthetic weight.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3054958,
              "author_name": "#1BuBu",
              "author_url": "",
              "post_date": "2024-11-25T11:23:57.230000",
              "content": "<p>The first solution is use statistical model to analyze it, as paper \"Multi-dimensional latent group structures with heterogeneous distributions\" here: <a href=\"https://www.sciencedirect.com/science/article/pii/S0304407621002177\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S0304407621002177</a>. Or we can just use Transformer to capture it automatically as paper \"Transformers as Statisticians: Provable In-Context<br>\nLearning with In-Context Algorithm Selection\" said here: <a href=\"https://proceedings.neurips.cc/paper_files/paper/2023/file/b2e63e36c57e153b9015fece2352a9f9-Paper-Conference.pdf\" target=\"_blank\">https://proceedings.neurips.cc/paper_files/paper/2023/file/b2e63e36c57e153b9015fece2352a9f9-Paper-Conference.pdf</a> 😂</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3086912,
          "author_name": "skrrydg",
          "author_url": "",
          "post_date": "2025-01-02T20:39:06.310000",
          "content": "<p>All features and target are stationary</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3054590,
      "author_name": "Param2007",
      "author_url": "",
      "post_date": "2024-11-24T21:10:25.540000",
      "content": "<p>Oh, but I had a quick question?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3054646,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-25T00:02:06.530000",
          "content": "<p>I think it's better ask it then wasting a chance to get the answer :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3086743,
      "author_name": "yb",
      "author_url": "",
      "post_date": "2025-01-02T16:32:11.910000",
      "content": "<p>Could you please provide insights or strategies for reducing the discrepancy between my local CV scores and the leaderboard (LB) scores? My main challenge seems to be that my local CV setup does not reliably indicate which models will perform better on the LB. Any advice on improving CV strategies or diagnosing this issue would be greatly appreciated.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3086718,
      "author_name": "Josh",
      "author_url": "",
      "post_date": "2025-01-02T16:10:16.347000",
      "content": "<p>new feature idea:  considering that the total sum of weights is stable each day across each time_id, do you think it's worth calculating for each time_id,  the weight as a % of weights sum ? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3056111,
      "author_name": "Michael Timbs",
      "author_url": "",
      "post_date": "2024-11-26T15:11:31.977000",
      "content": "<p>I finally got my first ensemble trained and submitted only to drop from 0.0034 baseline to -0.006 with the ensemble so that was fun!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3053871,
      "author_name": "Dennis",
      "author_url": "",
      "post_date": "2024-11-24T04:03:58.063000",
      "content": "<p>Can you talk about the use of tags in this competition?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3054304,
          "author_name": "benchouikha01",
          "author_url": "",
          "post_date": "2024-11-24T14:47:31.350000",
          "content": "<p>I tried to analyze the tags closely and I noticed that these are used to group the features (same for the responders tags) into features (or responders) that have the same behavior/characteristics. <br>\nAlso from what I have understood, the tags for the features have nothing to do with the tags for the responders.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3054509,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-24T18:42:20.010000",
              "content": "<p>Agree. On top that - tag is some sort of a constant, right? And since tags are the same for every single time step - I'd presume that neural network is capable of learning this kind of representation on it's own (if it's a useful one). So, I don't think it worth spending time and effort solving the \"tags quiz\", better to focus on a meaningful \"moving\" components instead… </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3054533,
              "author_name": "benchouikha01",
              "author_url": "",
              "post_date": "2024-11-24T19:27:33.347000",
              "content": "<p>Fully agree with you.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3053809,
      "author_name": "Ayman Allawi",
      "author_url": "",
      "post_date": "2024-11-24T00:03:11.847000",
      "content": "<p>Have you done any feature engineering?<br>\nand by feature engineering I mean adding new features from the existing one </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3053811,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-24T00:09:29.147000",
          "content": "<p>No, I think this is a dead end. The neural network would likely outperform any of us in \"feature engineering.\"<br>\nAnd in this particular domain feature engineering is twice as useless. It's kind of \"silver bullet\" story - to come up with an artificial feature that would serve as a predictor for the market…</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3053813,
              "author_name": "Ayman Allawi",
              "author_url": "",
              "post_date": "2024-11-24T00:18:38.920000",
              "content": "<p>In some time series problems feature engineering can add great value but sure no silver bullet you need to add good set of features <br>\nMaybe this competition is different on that regard for some reasons such as:<br>\n1- the nature of finicial market <br>\n2- some features are maybe already engineered features from other features that might or mightn't exist in the dataset <br>\nSo if you decide to do feature engineering you will do it totally blind </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3053277,
      "author_name": "I2nfinit3y",
      "author_url": "",
      "post_date": "2024-11-23T11:28:56.140000",
      "content": "<p>Do you think we should make different preprocessing for different symbol_id data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3053315,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-23T11:52:51.700000",
          "content": "<p>I don't know, why would you do that? In finance, we have a kind of a uniform scales - returns, prices, etc. You should have a strong reason to put symbols onto different scales, I guess…</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3054926,
          "author_name": "#1BuBu",
          "author_url": "",
          "post_date": "2024-11-25T10:44:01.643000",
          "content": "<p>To utilize symbol_id's heterogeneity, the submission will be timeout😂 and it's possible to add new features in test data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3053238,
      "author_name": "Sergei Fironov",
      "author_url": "",
      "post_date": "2024-11-23T10:07:28.343000",
      "content": "<p>Let's talk about normalization. Is there a good normalization for time series?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3053241,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-23T10:16:19.963000",
          "content": "<p>I would add \"non stationary time series\" on top of that :)</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3053391,
              "author_name": "Fernando Melo",
              "author_url": "",
              "post_date": "2024-11-23T12:42:55.457000",
              "content": "<p>Any tips on how to deal with non-stationary time series?<br>\nLSTM? GRU? Scaling?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3053796,
              "author_name": "Victor Shlepov",
              "author_url": "",
              "post_date": "2024-11-23T23:38:09.973000",
              "content": "<p>Look, without getting into unnecessary details- there's no \"magic bullet\" here. I'd rather put all this fancy acronyms aside and try to think in some real, industry-specific terms - what and why you want the model to mimic. That's exactly what I do, cause tweaking the code or juggling the different layers with no big picture in mind is no fun at all…</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3054300,
              "author_name": "Fernando Melo",
              "author_url": "",
              "post_date": "2024-11-24T14:43:27.553000",
              "content": "<p>Thx.<br>\nDo you have any materials to better understand this problem and “industry-specific terms”?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3059652,
              "author_name": "xuwuu",
              "author_url": "",
              "post_date": "2024-12-01T00:34:56.370000",
              "content": "<p>Hi, thank you so much for your insightful response! I'm not from the finance industry, so I’m not very familiar with the specific use cases in this domain. Could you provide an example or recommend some resources to help me better understand how to build models in finance? Specifically, how do you select appropriate features and understand the applications of different models for market prediction?</p>\n<p>Looking forward to your suggestions, and thanks again!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3052821,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2024-11-22T21:17:02.133000",
      "content": "<p>How about the topic is whichever of those was the driver behind you climbing to 3rd place so fast? 😝</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3053087,
          "author_name": "Victor Shlepov",
          "author_url": "",
          "post_date": "2024-11-23T06:08:33.713000",
          "content": "<p>I just stick to my own advice and keep circling back to the same ideas over and over again… :)</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3054748,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-25T05:16:36.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3054299,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-24T14:41:55.587000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3052807": "Hey folks, it was a good week, right? The top 10 is getting denser now, which is actually a positive sign - the more competition, the better! Let’s choose a topic for next week to discuss. No secret recipes - keep those to yourself (I don’t believe in any for the financial markets anyway) - just some high-level observations, insights, and ideas. Let’s make this competition a little more interactive and live!\nSo, what’s your call? Normalization, online learning, lags, augmentation strategy, or something else? Just drop your suggestions here. We’ll pick one, and I’ll share my thoughts; I hope you’ll do the same too!",
    "3052880": "Normalization. online learning in that order.",
    "3055706": "@victorshlepov can u share a simple baseline of online learning (NN) ? no matter how simple it is, I tried to build one, but got tons of bug😭",
    "3054940": "Non-stationary property is the key to this competition, as I noticed. The most important issue is, how to **fit a general model** and to **adjust it flexibly** according to the different pattern. The big problem is, we are limited to the time constraint and submission form in batch instead of time series of api. How to track this problem? 1. Form the data as data streaming with feature engineering or 2. form the data as time series with LSTM or Transformer to capture the non-stationary. Do you have any advice?",
    "3054590": "Oh, but I had a quick question?",
    "3086743": "Could you please provide insights or strategies for reducing the discrepancy between my local CV scores and the leaderboard (LB) scores? My main challenge seems to be that my local CV setup does not reliably indicate which models will perform better on the LB. Any advice on improving CV strategies or diagnosing this issue would be greatly appreciated.",
    "3086718": "new feature idea:  considering that the total sum of weights is stable each day across each time_id, do you think it's worth calculating for each time_id,  the weight as a % of weights sum ? \n\n\n",
    "3056111": "I finally got my first ensemble trained and submitted only to drop from 0.0034 baseline to -0.006 with the ensemble so that was fun!",
    "3053871": "Can you talk about the use of tags in this competition?",
    "3053809": "Have you done any feature engineering?\nand by feature engineering I mean adding new features from the existing one ",
    "3053277": "Do you think we should make different preprocessing for different symbol_id data?",
    "3053238": "Let's talk about normalization. Is there a good normalization for time series?",
    "3052821": "How about the topic is whichever of those was the driver behind you climbing to 3rd place so fast? 😝",
    "3054748": "",
    "3054299": ""
  }
}