{
  "id": 556470,
  "title": "What are your biggest learnings and frustrations from this competition?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/556470",
  "author_name": "",
  "post_date": "2025-01-13T13:40:54.704037200Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>In this competition, I faced three main problems that hindered me from achieving a better score:</p>\n<p><strong>Online Learning:</strong><br>\nI tried a lot, but I always ended up with worse results than my pure offline model. I believe my data pipeline is correct, but I am not sure if I implemented the correct parameters to retrain my neural network (NN). After this competition, I plan to set up an offline environment to learn more about this topic and check if someone shared an online learning notebook using NN.</p>\n<p><strong>Sequential Approach:</strong><br>\nOffline, my model worked well, but when I submitted the results, they worsened. This was probably due to a bug in my implementation logic, but I wasn’t able to figure out what went wrong.<br>\nAnother issue was that I spent a lot of time training my model with Keras and using ragged tensors. However, Kaggle notebooks use Keras 3, which doesn’t support ragged tensors. Installing another version of Keras/TensorFlow caused many problems. For the next competition, I plan to start with PyTorch or TensorFlow's Gradient Tape, as I believe these will better support online implementations.</p>\n<p><strong>Framework Challenges:</strong><br>\nAs mentioned earlier, I spent a significant amount of time working with Keras, where I faced issues with implementing ragged tensors and online learning. I know the problem lies with me, not the framework itself, but PyTorch seems more user-friendly for highly customized problems.</p>\n<p><strong>Lessons Learned:</strong></p>\n<p>I learned a lot of tweaks that can be applied when dealing with this type of problem, such as normalization and layer adjustments. Although I couldn’t achieve good results with sequential models, I now understand them better than I did before this competition. I will keep studying.</p>\n<p>My results so far are not great, but considering that I achieved almost 0.8 with an offline NN model, I believe that if I had successfully implemented online learning, my scores could have improved significantly.</p>",
  "messages": [
    {
      "id": "3095603",
      "postDate": "01/13/2025 13:40:54",
      "content": "<p>In this competition, I faced three main problems that hindered me from achieving a better score:</p>\n<p><strong>Online Learning:</strong><br>\nI tried a lot, but I always ended up with worse results than my pure offline model. I believe my data pipeline is correct, but I am not sure if I implemented the correct parameters to retrain my neural network (NN). After this competition, I plan to set up an offline environment to learn more about this topic and check if someone shared an online learning notebook using NN.</p>\n<p><strong>Sequential Approach:</strong><br>\nOffline, my model worked well, but when I submitted the results, they worsened. This was probably due to a bug in my implementation logic, but I wasn’t able to figure out what went wrong.<br>\nAnother issue was that I spent a lot of time training my model with Keras and using ragged tensors. However, Kaggle notebooks use Keras 3, which doesn’t support ragged tensors. Installing another version of Keras/TensorFlow caused many problems. For the next competition, I plan to start with PyTorch or TensorFlow's Gradient Tape, as I believe these will better support online implementations.</p>\n<p><strong>Framework Challenges:</strong><br>\nAs mentioned earlier, I spent a significant amount of time working with Keras, where I faced issues with implementing ragged tensors and online learning. I know the problem lies with me, not the framework itself, but PyTorch seems more user-friendly for highly customized problems.</p>\n<p><strong>Lessons Learned:</strong></p>\n<p>I learned a lot of tweaks that can be applied when dealing with this type of problem, such as normalization and layer adjustments. Although I couldn’t achieve good results with sequential models, I now understand them better than I did before this competition. I will keep studying.</p>\n<p>My results so far are not great, but considering that I achieved almost 0.8 with an offline NN model, I believe that if I had successfully implemented online learning, my scores could have improved significantly.</p>",
      "rawMarkdown": "In this competition, I faced three main problems that hindered me from achieving a better score:\n\n**Online Learning:**\nI tried a lot, but I always ended up with worse results than my pure offline model. I believe my data pipeline is correct, but I am not sure if I implemented the correct parameters to retrain my neural network (NN). After this competition, I plan to set up an offline environment to learn more about this topic and check if someone shared an online learning notebook using NN.\n\n**Sequential Approach:**\nOffline, my model worked well, but when I submitted the results, they worsened. This was probably due to a bug in my implementation logic, but I wasn’t able to figure out what went wrong.\nAnother issue was that I spent a lot of time training my model with Keras and using ragged tensors. However, Kaggle notebooks use Keras 3, which doesn’t support ragged tensors. Installing another version of Keras/TensorFlow caused many problems. For the next competition, I plan to start with PyTorch or TensorFlow's Gradient Tape, as I believe these will better support online implementations.\n\n**Framework Challenges:**\nAs mentioned earlier, I spent a significant amount of time working with Keras, where I faced issues with implementing ragged tensors and online learning. I know the problem lies with me, not the framework itself, but PyTorch seems more user-friendly for highly customized problems.\n\n**Lessons Learned:**\n\nI learned a lot of tweaks that can be applied when dealing with this type of problem, such as normalization and layer adjustments. Although I couldn’t achieve good results with sequential models, I now understand them better than I did before this competition. I will keep studying.\n\nMy results so far are not great, but considering that I achieved almost 0.8 with an offline NN model, I believe that if I had successfully implemented online learning, my scores could have improved significantly.",
      "votes": null
    },
    {
      "id": "3095645",
      "postDate": "01/13/2025 14:38:17",
      "content": "<p>AI tools / LLM's:<br>\nChat-GPT ,Claude and Notebook LM helped a lot to implement ideas, mainly for a student like me who has just done lot of theory but does not have that much practical knowledge.<br>\nThe learning curve was very fast with this tools, brainstorming, researching, validating the ideas, writing the skeleton code for it, etc.<br>\nIn this comp what I mostly learned is how to use this tools .<br>\nOur public score is not that good, but we were able to implement good online learning model which I hope in private evaluations will play a pivotal role.<br>\nImplementing models which actually improve with online learning is good I think ( makes me happy to see them work), considering I'm implementing it for the first time.</p>\n<p>I'm very curios as to what online learning techniques or even how online learning is implemented by top teams. </p>",
      "rawMarkdown": "AI tools / LLM's:\nChat-GPT ,Claude and Notebook LM helped a lot to implement ideas, mainly for a student like me who has just done lot of theory but does not have that much practical knowledge.\nThe learning curve was very fast with this tools, brainstorming, researching, validating the ideas, writing the skeleton code for it, etc.\nIn this comp what I mostly learned is how to use this tools .\nOur public score is not that good, but we were able to implement good online learning model which I hope in private evaluations will play a pivotal role.\nImplementing models which actually improve with online learning is good I think ( makes me happy to see them work), considering I'm implementing it for the first time.\n\nI'm very curios as to what online learning techniques or even how online learning is implemented by top teams.",
      "votes": null
    },
    {
      "id": "3095670",
      "postDate": "01/13/2025 15:15:43",
      "content": "<p>For me, the biggest take-away is testing things in the competition environment earlier. More than once I, I spent quite some time implementing some custom solution locally that I did not get running in the Kaggle notebook, either due to longer runtimes and thus runtime violations or difficulties importing manual modules into the notebook. </p>",
      "rawMarkdown": "For me, the biggest take-away is testing things in the competition environment earlier. More than once I, I spent quite some time implementing some custom solution locally that I did not get running in the Kaggle notebook, either due to longer runtimes and thus runtime violations or difficulties importing manual modules into the notebook.",
      "votes": null
    },
    {
      "id": "3095760",
      "postDate": "01/13/2025 17:13:00",
      "content": "<p>Me too. Half of my submissions failed because I did not do testing.</p>",
      "rawMarkdown": "Me too. Half of my submissions failed because I did not do testing.",
      "votes": null
    },
    {
      "id": "3095838",
      "postDate": "01/13/2025 20:05:30",
      "content": "<blockquote>\n  <p>My results so far are not great, but considering that I achieved almost 0.8 with an offline NN model, I believe that if I had successfully implemented online learning, my scores could have improved significantly.</p>\n</blockquote>\n<p>The competition isn't finished. You may have a good surprise in the private score. Public high-score solutions may not hold in the private dataset. Furthermore, as you said, you approached the problem from multiple angles and are leaving the competition with more knowledge than you had at the beginning. Let's keep doing it.</p>",
      "rawMarkdown": ">My results so far are not great, but considering that I achieved almost 0.8 with an offline NN model, I believe that if I had successfully implemented online learning, my scores could have improved significantly.\n\nThe competition isn't finished. You may have a good surprise in the private score. Public high-score solutions may not hold in the private dataset. Furthermore, as you said, you approached the problem from multiple angles and are leaving the competition with more knowledge than you had at the beginning. Let's keep doing it.",
      "votes": null
    },
    {
      "id": "3095844",
      "postDate": "01/13/2025 20:18:54",
      "content": "<p>The competition was very interesting and I could learn a lot. I can't wait to see the best solutions.</p>\n<p>If I could go back I would spend less time with the classic Gradient boosting trees methods and with feature engineering at least for me this did not allow me to advance much in the ranking.</p>\n<p>I would have preferred to work more on NN architecture and online learning. Only recently I managed to successfully implement an online learning of my NN however the result is less than my best solution based on GBDT, NN and a meta model that I update with online learning. Probably I need to refine the basic architecture.</p>\n<p>I had difficulty understanding how the API worked and the fact that at the beginning of each day ALL the lags of the previous day were provided but in the end thanks to the discussions and the simulators I was able to understand. I would have definitely spent more time creating a simulator to better test my models offline.</p>",
      "rawMarkdown": "The competition was very interesting and I could learn a lot. I can't wait to see the best solutions.\n\nIf I could go back I would spend less time with the classic Gradient boosting trees methods and with feature engineering at least for me this did not allow me to advance much in the ranking.\n\nI would have preferred to work more on NN architecture and online learning. Only recently I managed to successfully implement an online learning of my NN however the result is less than my best solution based on GBDT, NN and a meta model that I update with online learning. Probably I need to refine the basic architecture.\n\nI had difficulty understanding how the API worked and the fact that at the beginning of each day ALL the lags of the previous day were provided but in the end thanks to the discussions and the simulators I was able to understand. I would have definitely spent more time creating a simulator to better test my models offline.",
      "votes": null
    },
    {
      "id": "3095948",
      "postDate": "01/13/2025 23:55:52",
      "content": "<p>Biggest frustration- time series is really hard! I tried, and I failed miserably 🤣  <br>\nBiggest learning- hopefully did not happened yet and will happen when people in the ~top20 or so publish solutions. Please publish! 🤣</p>",
      "rawMarkdown": "Biggest frustration- time series is really hard! I tried, and I failed miserably 🤣  \nBiggest learning- hopefully did not happened yet and will happen when people in the ~top20 or so publish solutions. Please publish! 🤣",
      "votes": null
    },
    {
      "id": "3096022",
      "postDate": "01/14/2025 02:27:09",
      "content": "<p>I have been using some advanced time series models from the Darts package. The issue is that, it really needs the data timestamps to be aligned. It will throw exception even if your alignment is off by one. Suffice to say, having different number of <code>time_id</code>s in different dates, as well as ground truth responders being added at a different rate (daily), together with an opaque test server, makes the solution a whole lot harder than it needs to be.</p>",
      "rawMarkdown": "I have been using some advanced time series models from the Darts package. The issue is that, it really needs the data timestamps to be aligned. It will throw exception even if your alignment is off by one. Suffice to say, having different number of `time_id`s in different dates, as well as ground truth responders being added at a different rate (daily), together with an opaque test server, makes the solution a whole lot harder than it needs to be.",
      "votes": null
    },
    {
      "id": "3096035",
      "postDate": "01/14/2025 02:56:39",
      "content": "<p>Biggest frustration： Too late to succeed training online.</p>",
      "rawMarkdown": "Biggest frustration： Too late to succeed training online.",
      "votes": null
    },
    {
      "id": "3096040",
      "postDate": "01/14/2025 03:05:17",
      "content": "<p>The biggest purpose in this competition is to submit transformers model without failure, and then I succeeded it! I would like to aim to win with transformers next time in other competitions.</p>",
      "rawMarkdown": "The biggest purpose in this competition is to submit transformers model without failure, and then I succeeded it! I would like to aim to win with transformers next time in other competitions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3095645,
      "author_name": "zeroday19",
      "author_url": "",
      "post_date": "01/13/2025 14:38:17",
      "content": "<p>AI tools / LLM's:<br>\nChat-GPT ,Claude and Notebook LM helped a lot to implement ideas, mainly for a student like me who has just done lot of theory but does not have that much practical knowledge.<br>\nThe learning curve was very fast with this tools, brainstorming, researching, validating the ideas, writing the skeleton code for it, etc.<br>\nIn this comp what I mostly learned is how to use this tools .<br>\nOur public score is not that good, but we were able to implement good online learning model which I hope in private evaluations will play a pivotal role.<br>\nImplementing models which actually improve with online learning is good I think ( makes me happy to see them work), considering I'm implementing it for the first time.</p>\n<p>I'm very curios as to what online learning techniques or even how online learning is implemented by top teams. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3095670,
      "author_name": "jonassend",
      "author_url": "",
      "post_date": "01/13/2025 15:15:43",
      "content": "<p>For me, the biggest take-away is testing things in the competition environment earlier. More than once I, I spent quite some time implementing some custom solution locally that I did not get running in the Kaggle notebook, either due to longer runtimes and thus runtime violations or difficulties importing manual modules into the notebook. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3095760,
          "author_name": "ryanyutongwu",
          "author_url": "",
          "post_date": "01/13/2025 17:13:00",
          "content": "<p>Me too. Half of my submissions failed because I did not do testing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3095838,
      "author_name": "serjhenrique",
      "author_url": "",
      "post_date": "01/13/2025 20:05:30",
      "content": "<blockquote>\n  <p>My results so far are not great, but considering that I achieved almost 0.8 with an offline NN model, I believe that if I had successfully implemented online learning, my scores could have improved significantly.</p>\n</blockquote>\n<p>The competition isn't finished. You may have a good surprise in the private score. Public high-score solutions may not hold in the private dataset. Furthermore, as you said, you approached the problem from multiple angles and are leaving the competition with more knowledge than you had at the beginning. Let's keep doing it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3095844,
      "author_name": "simonedegasperis",
      "author_url": "",
      "post_date": "01/13/2025 20:18:54",
      "content": "<p>The competition was very interesting and I could learn a lot. I can't wait to see the best solutions.</p>\n<p>If I could go back I would spend less time with the classic Gradient boosting trees methods and with feature engineering at least for me this did not allow me to advance much in the ranking.</p>\n<p>I would have preferred to work more on NN architecture and online learning. Only recently I managed to successfully implement an online learning of my NN however the result is less than my best solution based on GBDT, NN and a meta model that I update with online learning. Probably I need to refine the basic architecture.</p>\n<p>I had difficulty understanding how the API worked and the fact that at the beginning of each day ALL the lags of the previous day were provided but in the end thanks to the discussions and the simulators I was able to understand. I would have definitely spent more time creating a simulator to better test my models offline.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3095948,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "01/13/2025 23:55:52",
      "content": "<p>Biggest frustration- time series is really hard! I tried, and I failed miserably 🤣  <br>\nBiggest learning- hopefully did not happened yet and will happen when people in the ~top20 or so publish solutions. Please publish! 🤣</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3096022,
      "author_name": "tinkei",
      "author_url": "",
      "post_date": "01/14/2025 02:27:09",
      "content": "<p>I have been using some advanced time series models from the Darts package. The issue is that, it really needs the data timestamps to be aligned. It will throw exception even if your alignment is off by one. Suffice to say, having different number of <code>time_id</code>s in different dates, as well as ground truth responders being added at a different rate (daily), together with an opaque test server, makes the solution a whole lot harder than it needs to be.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3096035,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "01/14/2025 02:56:39",
      "content": "<p>Biggest frustration： Too late to succeed training online.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3096040,
      "author_name": "hdynamics",
      "author_url": "",
      "post_date": "01/14/2025 03:05:17",
      "content": "<p>The biggest purpose in this competition is to submit transformers model without failure, and then I succeeded it! I would like to aim to win with transformers next time in other competitions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3095603": "In this competition, I faced three main problems that hindered me from achieving a better score:\n\n**Online Learning:**\nI tried a lot, but I always ended up with worse results than my pure offline model. I believe my data pipeline is correct, but I am not sure if I implemented the correct parameters to retrain my neural network (NN). After this competition, I plan to set up an offline environment to learn more about this topic and check if someone shared an online learning notebook using NN.\n\n**Sequential Approach:**\nOffline, my model worked well, but when I submitted the results, they worsened. This was probably due to a bug in my implementation logic, but I wasn’t able to figure out what went wrong.\nAnother issue was that I spent a lot of time training my model with Keras and using ragged tensors. However, Kaggle notebooks use Keras 3, which doesn’t support ragged tensors. Installing another version of Keras/TensorFlow caused many problems. For the next competition, I plan to start with PyTorch or TensorFlow's Gradient Tape, as I believe these will better support online implementations.\n\n**Framework Challenges:**\nAs mentioned earlier, I spent a significant amount of time working with Keras, where I faced issues with implementing ragged tensors and online learning. I know the problem lies with me, not the framework itself, but PyTorch seems more user-friendly for highly customized problems.\n\n**Lessons Learned:**\n\nI learned a lot of tweaks that can be applied when dealing with this type of problem, such as normalization and layer adjustments. Although I couldn’t achieve good results with sequential models, I now understand them better than I did before this competition. I will keep studying.\n\nMy results so far are not great, but considering that I achieved almost 0.8 with an offline NN model, I believe that if I had successfully implemented online learning, my scores could have improved significantly.",
    "3095645": "AI tools / LLM's:\nChat-GPT ,Claude and Notebook LM helped a lot to implement ideas, mainly for a student like me who has just done lot of theory but does not have that much practical knowledge.\nThe learning curve was very fast with this tools, brainstorming, researching, validating the ideas, writing the skeleton code for it, etc.\nIn this comp what I mostly learned is how to use this tools .\nOur public score is not that good, but we were able to implement good online learning model which I hope in private evaluations will play a pivotal role.\nImplementing models which actually improve with online learning is good I think ( makes me happy to see them work), considering I'm implementing it for the first time.\n\nI'm very curios as to what online learning techniques or even how online learning is implemented by top teams.",
    "3095670": "For me, the biggest take-away is testing things in the competition environment earlier. More than once I, I spent quite some time implementing some custom solution locally that I did not get running in the Kaggle notebook, either due to longer runtimes and thus runtime violations or difficulties importing manual modules into the notebook.",
    "3095760": "Me too. Half of my submissions failed because I did not do testing.",
    "3095838": ">My results so far are not great, but considering that I achieved almost 0.8 with an offline NN model, I believe that if I had successfully implemented online learning, my scores could have improved significantly.\n\nThe competition isn't finished. You may have a good surprise in the private score. Public high-score solutions may not hold in the private dataset. Furthermore, as you said, you approached the problem from multiple angles and are leaving the competition with more knowledge than you had at the beginning. Let's keep doing it.",
    "3095844": "The competition was very interesting and I could learn a lot. I can't wait to see the best solutions.\n\nIf I could go back I would spend less time with the classic Gradient boosting trees methods and with feature engineering at least for me this did not allow me to advance much in the ranking.\n\nI would have preferred to work more on NN architecture and online learning. Only recently I managed to successfully implement an online learning of my NN however the result is less than my best solution based on GBDT, NN and a meta model that I update with online learning. Probably I need to refine the basic architecture.\n\nI had difficulty understanding how the API worked and the fact that at the beginning of each day ALL the lags of the previous day were provided but in the end thanks to the discussions and the simulators I was able to understand. I would have definitely spent more time creating a simulator to better test my models offline.",
    "3095948": "Biggest frustration- time series is really hard! I tried, and I failed miserably 🤣  \nBiggest learning- hopefully did not happened yet and will happen when people in the ~top20 or so publish solutions. Please publish! 🤣",
    "3096022": "I have been using some advanced time series models from the Darts package. The issue is that, it really needs the data timestamps to be aligned. It will throw exception even if your alignment is off by one. Suffice to say, having different number of `time_id`s in different dates, as well as ground truth responders being added at a different rate (daily), together with an opaque test server, makes the solution a whole lot harder than it needs to be.",
    "3096035": "Biggest frustration： Too late to succeed training online.",
    "3096040": "The biggest purpose in this competition is to submit transformers model without failure, and then I succeeded it! I would like to aim to win with transformers next time in other competitions."
  },
  "source": "meta"
}