{
  "id": 556548,
  "title": "[Public LB 12th] Competition Wrap-up: Great Journey and Thank you all! ",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556548",
  "author_name": "SLi",
  "post_date": "2025-01-14T00:51:20.898000",
  "votes": 41,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Finally, it is time to wrap up this incredible journay with Jane Street and Kaggle. I have learned so much from this competition, and I am grateful for the opportunity to work on such challenging and rewarding projects. <br>\nI would like to express my gratitudes to the Jane Street and Kaggle for organizing this competition and providing us with the support. I would also like to thank other participants who have generously shared their knowledge and insights throughout this competition. My special thanks go to <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> , who have shared pure gold discussion posts that greatly enlightened me. </p>\n<h1>Some of the big learnings:</h1>\n<ol>\n<li><strong>Pipeline building</strong>: I realized in the very beginning that building a robust inference pipeline, especially with online learning, is crucial for this competition. However, it is definitly not trivial and requires lots of efforts on the code design and optimization. The evaluation api together with the hidden test set, which are not easy to hack, make it even more challenging. So I started with building a synthetic test set to help debugging the pipeline at the beginning, and shared this work in the community. This has helped me a lot and I am very happy to see it found to be useful by many people too (46 upvotes, 227 copies). </li>\n<li><strong>GBDT models vs Deep Learning</strong>: I stared with LightGBM and XGBoost as my baseline models, but I could not effectively improve their performance. So at very early stage, I have switched to neural networks, which have shown to be more powerful in this competition. I have tried different architectures, including MLP, GRU, and Transformer, as well as different training strategies. There were several weeks that I was stuck with negative R2 scores, during which I almost exhusted all possible model architetures that I know of (e.g. iTransformers, patchTST, Convnet, etc.). It was a very frustrating period, but finally I was lucky to find a working solution. I definitely benefited a lot from the public discussions, especially the ones from <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . These are pure gold and I would recommend everyone to read (and upvote) them. </li>\n<li><strong>Nature of the data</strong>: As the host described in the competition overview, we have encountered all kinds of challenges one can imagine. Fat tailed distributions, non-stationary time-series, noisy signals and so on. A good EDA, and especially some visulization pipelines can really help to build ituitive understanding of the data. My early EDA work was also shared in the community, but I felt it could be further improved. For instance, if I could have done an analysis on the responders as amazing as <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a> does in his great post, I would have figure out more ways to design the model. </li>\n<li><strong>Other tweaks</strong>: The frustrated try-and-error period was not really for nothing. I have greatly improved my understandings about many model archetectures in a practical way. These are the \"get hands dirty\" times that I did learn a lot. Some tweaks I learned including the different normalization methods, loss functions, feature fusion modules and training strategies. Although not all of them were useful in the end, I am happy to have tried them.</li>\n</ol>\n<h1>Short summary of my solution</h1>\n<p>I will keep this part short as there are 6 month remaining. </p>\n<ul>\n<li>For models, I designed two different architectures using basic ingredients including GRU, MLP and Transformer (symbol-wise attention). Under each architecture, feature maps were ensemble in two ways, resulting in 4 different models.</li>\n<li>Features are the 79 raw features excluding <code>9, 10, 11</code>, <code>time_id</code>, <code>weights</code>, as well as mean and std of the lagged responders. Missing values were filled with zeros.  All responders were used as targets (instead of only Responder_6). As <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> pointed out, using auxiliary targets can greatly boost both CV &amp; LB.</li>\n<li>Models were validated using the last 120 days, with both offline and online mode. Training sets includes three settings, i.e. 978 days, 800 days and 600 days. Eventually only models trained with the 978 and 800 days were used in the final ensemble (8 models).</li>\n<li><strong>Online learning</strong> was designed to update the model on a daily basis, using a similar setting as the training. Unlike <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> 's solution, I did not differentiate the responder_6 loss and auxiliary targets loss during the online update. The model updating is quite fast. It was about 0.5~0.7 sec per model per day. A full online training using every 120 or 200 days could further boost the score, however I did not implement it as it will complicate the whole pipeline quite a lot. The major concern is the 1-min limit. </li>\n</ul>\n<h1>Cheers!</h1>\n<p>At this moment, it is wayyy too early to say anything about the final ranking. The six month ahead will be the real challenge. I will keep my finger crossed and hope nothing in my pipeline breaks. I wish everyone good luck and gets the worthy rewards for the hard work. <br>\nThank you all!</p>",
  "messages": [
    {
      "id": 3095982,
      "postDate": "2025-01-14T00:51:20.900Z",
      "content": "<p>Finally, it is time to wrap up this incredible journay with Jane Street and Kaggle. I have learned so much from this competition, and I am grateful for the opportunity to work on such challenging and rewarding projects. <br>\nI would like to express my gratitudes to the Jane Street and Kaggle for organizing this competition and providing us with the support. I would also like to thank other participants who have generously shared their knowledge and insights throughout this competition. My special thanks go to <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> , who have shared pure gold discussion posts that greatly enlightened me. </p>\n<h1>Some of the big learnings:</h1>\n<ol>\n<li><strong>Pipeline building</strong>: I realized in the very beginning that building a robust inference pipeline, especially with online learning, is crucial for this competition. However, it is definitly not trivial and requires lots of efforts on the code design and optimization. The evaluation api together with the hidden test set, which are not easy to hack, make it even more challenging. So I started with building a synthetic test set to help debugging the pipeline at the beginning, and shared this work in the community. This has helped me a lot and I am very happy to see it found to be useful by many people too (46 upvotes, 227 copies). </li>\n<li><strong>GBDT models vs Deep Learning</strong>: I stared with LightGBM and XGBoost as my baseline models, but I could not effectively improve their performance. So at very early stage, I have switched to neural networks, which have shown to be more powerful in this competition. I have tried different architectures, including MLP, GRU, and Transformer, as well as different training strategies. There were several weeks that I was stuck with negative R2 scores, during which I almost exhusted all possible model architetures that I know of (e.g. iTransformers, patchTST, Convnet, etc.). It was a very frustrating period, but finally I was lucky to find a working solution. I definitely benefited a lot from the public discussions, especially the ones from <a href=\"https://www.kaggle.com/victorshlepov\" target=\"_blank\">@victorshlepov</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> . These are pure gold and I would recommend everyone to read (and upvote) them. </li>\n<li><strong>Nature of the data</strong>: As the host described in the competition overview, we have encountered all kinds of challenges one can imagine. Fat tailed distributions, non-stationary time-series, noisy signals and so on. A good EDA, and especially some visulization pipelines can really help to build ituitive understanding of the data. My early EDA work was also shared in the community, but I felt it could be further improved. For instance, if I could have done an analysis on the responders as amazing as <a href=\"https://www.kaggle.com/johnpayne0\" target=\"_blank\">@johnpayne0</a> does in his great post, I would have figure out more ways to design the model. </li>\n<li><strong>Other tweaks</strong>: The frustrated try-and-error period was not really for nothing. I have greatly improved my understandings about many model archetectures in a practical way. These are the \"get hands dirty\" times that I did learn a lot. Some tweaks I learned including the different normalization methods, loss functions, feature fusion modules and training strategies. Although not all of them were useful in the end, I am happy to have tried them.</li>\n</ol>\n<h1>Short summary of my solution</h1>\n<p>I will keep this part short as there are 6 month remaining. </p>\n<ul>\n<li>For models, I designed two different architectures using basic ingredients including GRU, MLP and Transformer (symbol-wise attention). Under each architecture, feature maps were ensemble in two ways, resulting in 4 different models.</li>\n<li>Features are the 79 raw features excluding <code>9, 10, 11</code>, <code>time_id</code>, <code>weights</code>, as well as mean and std of the lagged responders. Missing values were filled with zeros.  All responders were used as targets (instead of only Responder_6). As <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> pointed out, using auxiliary targets can greatly boost both CV &amp; LB.</li>\n<li>Models were validated using the last 120 days, with both offline and online mode. Training sets includes three settings, i.e. 978 days, 800 days and 600 days. Eventually only models trained with the 978 and 800 days were used in the final ensemble (8 models).</li>\n<li><strong>Online learning</strong> was designed to update the model on a daily basis, using a similar setting as the training. Unlike <a href=\"https://www.kaggle.com/eivolkova\" target=\"_blank\">@eivolkova</a> 's solution, I did not differentiate the responder_6 loss and auxiliary targets loss during the online update. The model updating is quite fast. It was about 0.5~0.7 sec per model per day. A full online training using every 120 or 200 days could further boost the score, however I did not implement it as it will complicate the whole pipeline quite a lot. The major concern is the 1-min limit. </li>\n</ul>\n<h1>Cheers!</h1>\n<p>At this moment, it is wayyy too early to say anything about the final ranking. The six month ahead will be the real challenge. I will keep my finger crossed and hope nothing in my pipeline breaks. I wish everyone good luck and gets the worthy rewards for the hard work. <br>\nThank you all!</p>",
      "rawMarkdown": "\nFinally, it is time to wrap up this incredible journay with Jane Street and Kaggle. I have learned so much from this competition, and I am grateful for the opportunity to work on such challenging and rewarding projects. \n\nI would like to express my gratitudes to the Jane Street and Kaggle for organizing this competition and providing us with the support. I would also like to thank other participants who have generously shared their knowledge and insights throughout this competition. My special thanks go to @victorshlepov and @lihaorocky , who have shared pure gold discussion posts that greatly enlightened me. \n\n# Some of the big learnings:\n\n1. **Pipeline building**: I realized in the very beginning that building a robust inference pipeline, especially with online learning, is crucial for this competition. However, it is definitly not trivial and requires lots of efforts on the code design and optimization. The evaluation api together with the hidden test set, which are not easy to hack, make it even more challenging. So I started with building a synthetic test set to help debugging the pipeline at the beginning, and shared this work in the community. This has helped me a lot and I am very happy to see it found to be useful by many people too (46 upvotes, 227 copies). \n\n2. **GBDT models vs Deep Learning**: I stared with LightGBM and XGBoost as my baseline models, but I could not effectively improve their performance. So at very early stage, I have switched to neural networks, which have shown to be more powerful in this competition. I have tried different architectures, including MLP, GRU, and Transformer, as well as different training strategies. There were several weeks that I was stuck with negative R2 scores, during which I almost exhusted all possible model architetures that I know of (e.g. iTransformers, patchTST, Convnet, etc.). It was a very frustrating period, but finally I was lucky to find a working solution. I definitely benefited a lot from the public discussions, especially the ones from @victorshlepov and @lihaorocky . These are pure gold and I would recommend everyone to read (and upvote) them. \n\n3. **Nature of the data**: As the host described in the competition overview, we have encountered all kinds of challenges one can imagine. Fat tailed distributions, non-stationary time-series, noisy signals and so on. A good EDA, and especially some visulization pipelines can really help to build ituitive understanding of the data. My early EDA work was also shared in the community, but I felt it could be further improved. For instance, if I could have done an analysis on the responders as amazing as @johnpayne0 does in his great post, I would have figure out more ways to design the model. \n\n4. **Other tweaks**: The frustrated try-and-error period was not really for nothing. I have greatly improved my understandings about many model archetectures in a practical way. These are the \"get hands dirty\" times that I did learn a lot. Some tweaks I learned including the different normalization methods, loss functions, feature fusion modules and training strategies. Although not all of them were useful in the end, I am happy to have tried them.\n\n\n# Short summary of my solution\n\nI will keep this part short as there are 6 month remaining. \n\n* For models, I designed two different architectures using basic ingredients including GRU, MLP and Transformer (symbol-wise attention). Under each architecture, feature maps were ensemble in two ways, resulting in 4 different models.\n\n* Features are the 79 raw features excluding `9, 10, 11`, `time_id`, `weights`, as well as mean and std of the lagged responders. Missing values were filled with zeros.  All responders were used as targets (instead of only Responder_6). As @eivolkova pointed out, using auxiliary targets can greatly boost both CV & LB.\n\n* Models were validated using the last 120 days, with both offline and online mode. Training sets includes three settings, i.e. 978 days, 800 days and 600 days. Eventually only models trained with the 978 and 800 days were used in the final ensemble (8 models).\n\n* **Online learning** was designed to update the model on a daily basis, using a similar setting as the training. Unlike @eivolkova 's solution, I did not differentiate the responder_6 loss and auxiliary targets loss during the online update. The model updating is quite fast. It was about 0.5~0.7 sec per model per day. A full online training using every 120 or 200 days could further boost the score, however I did not implement it as it will complicate the whole pipeline quite a lot. The major concern is the 1-min limit. \n\n# Cheers!\n\nAt this moment, it is wayyy too early to say anything about the final ranking. The six month ahead will be the real challenge. I will keep my finger crossed and hope nothing in my pipeline breaks. I wish everyone good luck and gets the worthy rewards for the hard work. \n\n\nThank you all!",
      "votes": 41
    },
    {
      "id": 3096432,
      "postDate": "2025-01-14T11:46:01.717Z",
      "content": "<p>Thanks for the mentioning. It's fancinating to see your progress in this competition especially during the last 4 weeks. Great work and open discussion (guess this is so called kaggle spirit😀). </p>",
      "rawMarkdown": "Thanks for the mentioning. It's fancinating to see your progress in this competition especially during the last 4 weeks. Great work and open discussion (guess this is so called kaggle spirit😀). ",
      "votes": 3
    },
    {
      "id": 3100872,
      "postDate": "2025-01-20T03:31:52.143Z",
      "content": "<p>Thanks for your write up. During the competition, the suggestions you provided were helpful and inspiring. Thank you.<br>\nWould you consider publish your NN model structures in later times?</p>",
      "rawMarkdown": "Thanks for your write up. During the competition, the suggestions you provided were helpful and inspiring. Thank you.\nWould you consider publish your NN model structures in later times?"
    },
    {
      "id": 3096231,
      "postDate": "2025-01-14T06:40:13.343Z",
      "content": "<p>Thanks Sli for sharing your solutions. You submission helped me a lot during the competition. </p>\n<p>I used mlp to train the whole dataset, with global normalized features as ['time_id', 'symbol_id' , 79 features, 8 lags], hidden_sizes as [2048, 1024, 512, 256, 128, 64]  and dropout_rate = 0.6. With online learning using pytorch to learn row by row. I cannot get my lb score more than 0.0070 for this single model. What else did I miss here compared to yours, how can i improve. Thanks.</p>",
      "rawMarkdown": "Thanks Sli for sharing your solutions. You submission helped me a lot during the competition. \n\nI used mlp to train the whole dataset, with global normalized features as ['time_id', 'symbol_id' , 79 features, 8 lags], hidden_sizes as [2048, 1024, 512, 256, 128, 64]  and dropout_rate = 0.6. With online learning using pytorch to learn row by row. I cannot get my lb score more than 0.0070 for this single model. What else did I miss here compared to yours, how can i improve. Thanks."
    },
    {
      "id": 3096198,
      "postDate": "2025-01-14T06:19:28.620Z",
      "content": "<p>I've tried transformer with symbol-wise attention too, but I couldn't get good results and gave up too easily. Thanks for sharing.</p>",
      "rawMarkdown": "I've tried transformer with symbol-wise attention too, but I couldn't get good results and gave up too easily. Thanks for sharing."
    },
    {
      "id": 3096641,
      "postDate": "2025-01-14T14:58:54.090Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3095990,
      "postDate": "2025-01-14T00:59:03.530Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3095989,
      "postDate": "2025-01-14T00:57:57.540Z",
      "content": "<p>Thank you for sharing! </p>",
      "rawMarkdown": "Thank you for sharing! "
    }
  ],
  "comments": [
    {
      "id": 3096432,
      "author_name": "HAO",
      "author_url": "",
      "post_date": "2025-01-14T11:46:01.717000",
      "content": "<p>Thanks for the mentioning. It's fancinating to see your progress in this competition especially during the last 4 weeks. Great work and open discussion (guess this is so called kaggle spirit😀). </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3100872,
      "author_name": "ZT",
      "author_url": "",
      "post_date": "2025-01-20T03:31:52.143000",
      "content": "<p>Thanks for your write up. During the competition, the suggestions you provided were helpful and inspiring. Thank you.<br>\nWould you consider publish your NN model structures in later times?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096231,
      "author_name": "Yimin Hu",
      "author_url": "",
      "post_date": "2025-01-14T06:40:13.343000",
      "content": "<p>Thanks Sli for sharing your solutions. You submission helped me a lot during the competition. </p>\n<p>I used mlp to train the whole dataset, with global normalized features as ['time_id', 'symbol_id' , 79 features, 8 lags], hidden_sizes as [2048, 1024, 512, 256, 128, 64]  and dropout_rate = 0.6. With online learning using pytorch to learn row by row. I cannot get my lb score more than 0.0070 for this single model. What else did I miss here compared to yours, how can i improve. Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096198,
      "author_name": "CodeHacker",
      "author_url": "",
      "post_date": "2025-01-14T06:19:28.620000",
      "content": "<p>I've tried transformer with symbol-wise attention too, but I couldn't get good results and gave up too easily. Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3096641,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T14:58:54.090000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3095990,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-14T00:59:03.530000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3095989,
      "author_name": "鸽鸽257",
      "author_url": "",
      "post_date": "2025-01-14T00:57:57.540000",
      "content": "<p>Thank you for sharing! </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3095982": "\nFinally, it is time to wrap up this incredible journay with Jane Street and Kaggle. I have learned so much from this competition, and I am grateful for the opportunity to work on such challenging and rewarding projects. \n\nI would like to express my gratitudes to the Jane Street and Kaggle for organizing this competition and providing us with the support. I would also like to thank other participants who have generously shared their knowledge and insights throughout this competition. My special thanks go to @victorshlepov and @lihaorocky , who have shared pure gold discussion posts that greatly enlightened me. \n\n# Some of the big learnings:\n\n1. **Pipeline building**: I realized in the very beginning that building a robust inference pipeline, especially with online learning, is crucial for this competition. However, it is definitly not trivial and requires lots of efforts on the code design and optimization. The evaluation api together with the hidden test set, which are not easy to hack, make it even more challenging. So I started with building a synthetic test set to help debugging the pipeline at the beginning, and shared this work in the community. This has helped me a lot and I am very happy to see it found to be useful by many people too (46 upvotes, 227 copies). \n\n2. **GBDT models vs Deep Learning**: I stared with LightGBM and XGBoost as my baseline models, but I could not effectively improve their performance. So at very early stage, I have switched to neural networks, which have shown to be more powerful in this competition. I have tried different architectures, including MLP, GRU, and Transformer, as well as different training strategies. There were several weeks that I was stuck with negative R2 scores, during which I almost exhusted all possible model architetures that I know of (e.g. iTransformers, patchTST, Convnet, etc.). It was a very frustrating period, but finally I was lucky to find a working solution. I definitely benefited a lot from the public discussions, especially the ones from @victorshlepov and @lihaorocky . These are pure gold and I would recommend everyone to read (and upvote) them. \n\n3. **Nature of the data**: As the host described in the competition overview, we have encountered all kinds of challenges one can imagine. Fat tailed distributions, non-stationary time-series, noisy signals and so on. A good EDA, and especially some visulization pipelines can really help to build ituitive understanding of the data. My early EDA work was also shared in the community, but I felt it could be further improved. For instance, if I could have done an analysis on the responders as amazing as @johnpayne0 does in his great post, I would have figure out more ways to design the model. \n\n4. **Other tweaks**: The frustrated try-and-error period was not really for nothing. I have greatly improved my understandings about many model archetectures in a practical way. These are the \"get hands dirty\" times that I did learn a lot. Some tweaks I learned including the different normalization methods, loss functions, feature fusion modules and training strategies. Although not all of them were useful in the end, I am happy to have tried them.\n\n\n# Short summary of my solution\n\nI will keep this part short as there are 6 month remaining. \n\n* For models, I designed two different architectures using basic ingredients including GRU, MLP and Transformer (symbol-wise attention). Under each architecture, feature maps were ensemble in two ways, resulting in 4 different models.\n\n* Features are the 79 raw features excluding `9, 10, 11`, `time_id`, `weights`, as well as mean and std of the lagged responders. Missing values were filled with zeros.  All responders were used as targets (instead of only Responder_6). As @eivolkova pointed out, using auxiliary targets can greatly boost both CV & LB.\n\n* Models were validated using the last 120 days, with both offline and online mode. Training sets includes three settings, i.e. 978 days, 800 days and 600 days. Eventually only models trained with the 978 and 800 days were used in the final ensemble (8 models).\n\n* **Online learning** was designed to update the model on a daily basis, using a similar setting as the training. Unlike @eivolkova 's solution, I did not differentiate the responder_6 loss and auxiliary targets loss during the online update. The model updating is quite fast. It was about 0.5~0.7 sec per model per day. A full online training using every 120 or 200 days could further boost the score, however I did not implement it as it will complicate the whole pipeline quite a lot. The major concern is the 1-min limit. \n\n# Cheers!\n\nAt this moment, it is wayyy too early to say anything about the final ranking. The six month ahead will be the real challenge. I will keep my finger crossed and hope nothing in my pipeline breaks. I wish everyone good luck and gets the worthy rewards for the hard work. \n\n\nThank you all!",
    "3096432": "Thanks for the mentioning. It's fancinating to see your progress in this competition especially during the last 4 weeks. Great work and open discussion (guess this is so called kaggle spirit😀). ",
    "3100872": "Thanks for your write up. During the competition, the suggestions you provided were helpful and inspiring. Thank you.\nWould you consider publish your NN model structures in later times?",
    "3096231": "Thanks Sli for sharing your solutions. You submission helped me a lot during the competition. \n\nI used mlp to train the whole dataset, with global normalized features as ['time_id', 'symbol_id' , 79 features, 8 lags], hidden_sizes as [2048, 1024, 512, 256, 128, 64]  and dropout_rate = 0.6. With online learning using pytorch to learn row by row. I cannot get my lb score more than 0.0070 for this single model. What else did I miss here compared to yours, how can i improve. Thanks.",
    "3096198": "I've tried transformer with symbol-wise attention too, but I couldn't get good results and gave up too easily. Thanks for sharing.",
    "3096641": "",
    "3095990": "",
    "3095989": "Thank you for sharing! "
  }
}