{
  "id": 541396,
  "title": "Any success submitting with tensorflow : Need Help",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/541396",
  "author_name": "",
  "post_date": "2024-10-19T08:28:33.253038400Z",
  "votes": 8,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I'm trying to submit a very light neural network by it timed out. I have no clue of why it didn't work.</p>\n<p>I will take any insight if you were successful while submitting with tensorflow or not.</p>\n<p><strong>Related Notebook</strong>: <a href=\"https://www.kaggle.com/code/ulrich07/js-tensorflow-debug/notebook?scriptVersionId=202289829\" target=\"_blank\">link</a></p>\n<p>Happy Kaggling</p>\n<p><strong>Update</strong>: I found out why it timed out. A single model took ~3 hours to be scored. Since I used 5 fold-models it timed out. But it is very strange that a very light model takes  ~3 hours to be scored.</p>",
  "messages": [
    {
      "id": "3022073",
      "postDate": "10/19/2024 08:28:33",
      "content": "<p>Hello everyone,</p>\n<p>I'm trying to submit a very light neural network by it timed out. I have no clue of why it didn't work.</p>\n<p>I will take any insight if you were successful while submitting with tensorflow or not.</p>\n<p><strong>Related Notebook</strong>: <a href=\"https://www.kaggle.com/code/ulrich07/js-tensorflow-debug/notebook?scriptVersionId=202289829\" target=\"_blank\">link</a></p>\n<p>Happy Kaggling</p>\n<p><strong>Update</strong>: I found out why it timed out. A single model took ~3 hours to be scored. Since I used 5 fold-models it timed out. But it is very strange that a very light model takes  ~3 hours to be scored.</p>",
      "rawMarkdown": "Hello everyone,\n\nI'm trying to submit a very light neural network by it timed out. I have no clue of why it didn't work.\n\nI will take any insight if you were successful while submitting with tensorflow or not.\n\n**Related Notebook**: [link](https://www.kaggle.com/code/ulrich07/js-tensorflow-debug/notebook?scriptVersionId=202289829)\n\nHappy Kaggling\n\n**Update**: I found out why it timed out. A single model took ~3 hours to be scored. Since I used 5 fold-models it timed out. But it is very strange that a very light model takes  ~3 hours to be scored.",
      "votes": null
    },
    {
      "id": "3022334",
      "postDate": "10/19/2024 13:15:24",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> Just to confirm, are you following the rules below?</p>\n<pre><code>Each batch  predictions (except  very ) must be returned   minutes   batch features being provided.\n</code></pre>\n<pre><code>When your notebook is   the hidden  , inference_server.serve must be called within 15 minutes of the notebook starting or the gateway will throw  .  you need  than 15 minutes to load your model you can   during the very first  call,  does not have the usual 10 minute response deadline.\n</code></pre>",
      "rawMarkdown": "ulrich07 Just to confirm, are you following the rules below?\n\n~~~\nEach batch of predictions (except the very first) must be returned within 10 minutes of the batch features being provided.\n~~~\n\n~~~\nWhen your notebook is run on the hidden test set, inference_server.serve must be called within 15 minutes of the notebook starting or the gateway will throw an error. If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 10 minute response deadline.\n~~~",
      "votes": null
    },
    {
      "id": "3022446",
      "postDate": "10/19/2024 15:24:02",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a>. The model is loaded in 5 seconds, and can predict 2M rows in 30 seconds. I will share a notebook.</p>",
      "rawMarkdown": "Thank you @chumajin. The model is loaded in 5 seconds, and can predict 2M rows in 30 seconds. I will share a notebook.",
      "votes": null
    },
    {
      "id": "3022590",
      "postDate": "10/19/2024 18:14:07",
      "content": "<p>Are you using lags? Also facing a simple Ridge model timing out.. </p>",
      "rawMarkdown": "Are you using lags? Also facing a simple Ridge model timing out..",
      "votes": null
    },
    {
      "id": "3022669",
      "postDate": "10/19/2024 20:13:38",
      "content": "<p>Hiya, can you use your GPU allowance? I have definitely verified my account (hence I can take part in this competition) but turning on an accelerator mode doesn't work with the likes of XGBoost with cuda enabled or LGB…   </p>",
      "rawMarkdown": "Hiya, can you use your GPU allowance? I have definitely verified my account (hence I can take part in this competition) but turning on an accelerator mode doesn't work with the likes of XGBoost with cuda enabled or LGB...",
      "votes": null
    },
    {
      "id": "3022776",
      "postDate": "10/19/2024 23:48:08",
      "content": "<p>I don't think lags is the problem. I have a successfull rige (with and without lags). My TF model do not use lags.</p>",
      "rawMarkdown": "I don't think lags is the problem. I have a successfull rige (with and without lags). My TF model do not use lags.",
      "votes": null
    },
    {
      "id": "3022893",
      "postDate": "10/20/2024 03:48:11",
      "content": "<p>try smaller batch size and see how long does it take for your model to process<br>\nyour prediction callback is called on a ~30 batch size for 4mm rows</p>",
      "rawMarkdown": "try smaller batch size and see how long does it take for your model to process\nyour prediction callback is called on a ~30 batch size for 4mm rows",
      "votes": null
    },
    {
      "id": "3023316",
      "postDate": "10/20/2024 12:48:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a>  I add the notebook  link, can you have a look at it ? Thks in advance.</p>",
      "rawMarkdown": "Hi @chumajin  I add the notebook  link, can you have a look at it ? Thks in advance.",
      "votes": null
    },
    {
      "id": "3023360",
      "postDate": "10/20/2024 13:30:17",
      "content": "<p>Did you try removing batch size argument from predict? I ran a simple keras model at one point and didn’t have any issue.</p>",
      "rawMarkdown": "Did you try removing batch size argument from predict? I ran a simple keras model at one point and didn’t have any issue.",
      "votes": null
    },
    {
      "id": "3023361",
      "postDate": "10/20/2024 13:32:32",
      "content": "<p>Many thanks <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>. I will remove the batch_size parameter.</p>\n<p><strong>Update</strong>: But It's like what you said doesn't work for me. I removed batch_size on a private notebook but I guess it is still not working (submission is still runing for more than 1 hour now). Can you please share a keras code snipet or even check my <a href=\"https://www.kaggle.com/code/ulrich07/js-tensorflow-debug\" target=\"_blank\">notebook</a>.</p>",
      "rawMarkdown": "Many thanks @simonveitner. I will remove the batch_size parameter.\n\n**Update**: But It's like what you said doesn't work for me. I removed batch_size on a private notebook but I guess it is still not working (submission is still runing for more than 1 hour now). Can you please share a keras code snipet or even check my [notebook](https://www.kaggle.com/code/ulrich07/js-tensorflow-debug).",
      "votes": null
    },
    {
      "id": "3023409",
      "postDate": "10/20/2024 14:18:08",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> OK. I can see it. I will check it !</p>",
      "rawMarkdown": "ulrich07 OK. I can see it. I will check it !",
      "votes": null
    },
    {
      "id": "3023435",
      "postDate": "10/20/2024 14:40:13",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> I created the simulator notebook <a href=\"https://www.kaggle.com/code/chumajin/janestreet-simulator-for-time-series-api\" target=\"_blank\">here</a>. </p>\n<p>Now, I'm trying your model, maybe it will take about 2 hours 30 mins.<br>\nSo far, I think your model is fine. Therefore, there might be a bug in the time series API or something error we don't know.</p>\n<p>But I'll run everything through and let you know the results tomorrow morning!</p>",
      "rawMarkdown": "ulrich07 I created the simulator notebook [here](https://www.kaggle.com/code/chumajin/janestreet-simulator-for-time-series-api). \n\nNow, I'm trying your model, maybe it will take about 2 hours 30 mins.\nSo far, I think your model is fine. Therefore, there might be a bug in the time series API or something error we don't know.\n\nBut I'll run everything through and let you know the results tomorrow morning!",
      "votes": null
    },
    {
      "id": "3023462",
      "postDate": "10/20/2024 15:09:42",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> could you report here if the model ran through? maybe it's just a time issue.</p>",
      "rawMarkdown": "ulrich07 could you report here if the model ran through? maybe it's just a time issue.",
      "votes": null
    },
    {
      "id": "3023499",
      "postDate": "10/20/2024 16:14:38",
      "content": "<p>Yes It ran. But didn't check the time probably in ~3 hours</p>",
      "rawMarkdown": "Yes It ran. But didn't check the time probably in ~3 hours",
      "votes": null
    },
    {
      "id": "3023580",
      "postDate": "10/20/2024 18:03:03",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> I also succeeded with your code(batch size 512). For reference, I changed like this in order to insurance.</p>\n<pre><code>  test.select(FE).to_pandas().fillna().values\n・・・\n  predictions.fill_nan()\n</code></pre>\n<p>and the score was -14.4161</p>",
      "rawMarkdown": "ulrich07 I also succeeded with your code(batch size 512). For reference, I changed like this in order to insurance.\n\n~~~\nx = test.select(FE).to_pandas().fillna(3).values\n・・・\npredictions = predictions.fill_nan(0)\n~~~\n\n and the score was -14.4161",
      "votes": null
    },
    {
      "id": "3023593",
      "postDate": "10/20/2024 18:13:52",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a>.  I got -26.xxx score. I am very sad because as it is, putting a TF neural net in the final blend may be problematic concerning timeout issues. But <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>  how long does it take to score your submission. Thks in advance.</p>",
      "rawMarkdown": "Thank you @chumajin.  I got -26.xxx score. I am very sad because as it is, putting a TF neural net in the final blend may be problematic concerning timeout issues. But @simonveitner  how long does it take to score your submission. Thks in advance.",
      "votes": null
    },
    {
      "id": "3023598",
      "postDate": "10/20/2024 18:16:38",
      "content": "<p>I think it took less</p>",
      "rawMarkdown": "I think it took less",
      "votes": null
    },
    {
      "id": "3023682",
      "postDate": "10/20/2024 20:36:22",
      "content": "<p>I ran your code and there was no problem. The public score is -16.xxxx and took ~4hrs. </p>\n<p>Here is my kernel: <a href=\"https://www.kaggle.com/code/shiyili/js-tensorflow-debug\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js-tensorflow-debug</a> </p>",
      "rawMarkdown": "I ran your code and there was no problem. The public score is -16.xxxx and took ~4hrs. \n\nHere is my kernel: https://www.kaggle.com/code/shiyili/js-tensorflow-debug",
      "votes": null
    },
    {
      "id": "3024018",
      "postDate": "10/21/2024 08:50:18",
      "content": "<p>Given the current API and amount of data, I feel it is very hard to do anything complicated within every batch. </p>",
      "rawMarkdown": "Given the current API and amount of data, I feel it is very hard to do anything complicated within every batch.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3022334,
      "author_name": "chumajin",
      "author_url": "",
      "post_date": "10/19/2024 13:15:24",
      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> Just to confirm, are you following the rules below?</p>\n<pre><code>Each batch  predictions (except  very ) must be returned   minutes   batch features being provided.\n</code></pre>\n<pre><code>When your notebook is   the hidden  , inference_server.serve must be called within 15 minutes of the notebook starting or the gateway will throw  .  you need  than 15 minutes to load your model you can   during the very first  call,  does not have the usual 10 minute response deadline.\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3022446,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "10/19/2024 15:24:02",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a>. The model is loaded in 5 seconds, and can predict 2M rows in 30 seconds. I will share a notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3023316,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "10/20/2024 12:48:40",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a>  I add the notebook  link, can you have a look at it ? Thks in advance.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3023360,
              "author_name": "simonveitner",
              "author_url": "",
              "post_date": "10/20/2024 13:30:17",
              "content": "<p>Did you try removing batch size argument from predict? I ran a simple keras model at one point and didn’t have any issue.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3023361,
                  "author_name": "ulrich07",
                  "author_url": "",
                  "post_date": "10/20/2024 13:32:32",
                  "content": "<p>Many thanks <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>. I will remove the batch_size parameter.</p>\n<p><strong>Update</strong>: But It's like what you said doesn't work for me. I removed batch_size on a private notebook but I guess it is still not working (submission is still runing for more than 1 hour now). Can you please share a keras code snipet or even check my <a href=\"https://www.kaggle.com/code/ulrich07/js-tensorflow-debug\" target=\"_blank\">notebook</a>.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3023409,
                      "author_name": "chumajin",
                      "author_url": "",
                      "post_date": "10/20/2024 14:18:08",
                      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> OK. I can see it. I will check it !</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3023435,
                          "author_name": "chumajin",
                          "author_url": "",
                          "post_date": "10/20/2024 14:40:13",
                          "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> I created the simulator notebook <a href=\"https://www.kaggle.com/code/chumajin/janestreet-simulator-for-time-series-api\" target=\"_blank\">here</a>. </p>\n<p>Now, I'm trying your model, maybe it will take about 2 hours 30 mins.<br>\nSo far, I think your model is fine. Therefore, there might be a bug in the time series API or something error we don't know.</p>\n<p>But I'll run everything through and let you know the results tomorrow morning!</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3023462,
                              "author_name": "simonveitner",
                              "author_url": "",
                              "post_date": "10/20/2024 15:09:42",
                              "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> could you report here if the model ran through? maybe it's just a time issue.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3023499,
                                  "author_name": "ulrich07",
                                  "author_url": "",
                                  "post_date": "10/20/2024 16:14:38",
                                  "content": "<p>Yes It ran. But didn't check the time probably in ~3 hours</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 3023580,
                                      "author_name": "chumajin",
                                      "author_url": "",
                                      "post_date": "10/20/2024 18:03:03",
                                      "content": "<p><a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a> I also succeeded with your code(batch size 512). For reference, I changed like this in order to insurance.</p>\n<pre><code>  test.select(FE).to_pandas().fillna().values\n・・・\n  predictions.fill_nan()\n</code></pre>\n<p>and the score was -14.4161</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 3023593,
                                          "author_name": "ulrich07",
                                          "author_url": "",
                                          "post_date": "10/20/2024 18:13:52",
                                          "content": "<p>Thank you <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a>.  I got -26.xxx score. I am very sad because as it is, putting a TF neural net in the final blend may be problematic concerning timeout issues. But <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>  how long does it take to score your submission. Thks in advance.</p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 3023598,
                                              "author_name": "simonveitner",
                                              "author_url": "",
                                              "post_date": "10/20/2024 18:16:38",
                                              "content": "<p>I think it took less</p>",
                                              "votes": null,
                                              "replies": []
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3022590,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "10/19/2024 18:14:07",
      "content": "<p>Are you using lags? Also facing a simple Ridge model timing out.. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3022776,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "10/19/2024 23:48:08",
          "content": "<p>I don't think lags is the problem. I have a successfull rige (with and without lags). My TF model do not use lags.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3022669,
      "author_name": "missgranger",
      "author_url": "",
      "post_date": "10/19/2024 20:13:38",
      "content": "<p>Hiya, can you use your GPU allowance? I have definitely verified my account (hence I can take part in this competition) but turning on an accelerator mode doesn't work with the likes of XGBoost with cuda enabled or LGB…   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3022893,
      "author_name": "dc260123",
      "author_url": "",
      "post_date": "10/20/2024 03:48:11",
      "content": "<p>try smaller batch size and see how long does it take for your model to process<br>\nyour prediction callback is called on a ~30 batch size for 4mm rows</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3023682,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/20/2024 20:36:22",
      "content": "<p>I ran your code and there was no problem. The public score is -16.xxxx and took ~4hrs. </p>\n<p>Here is my kernel: <a href=\"https://www.kaggle.com/code/shiyili/js-tensorflow-debug\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js-tensorflow-debug</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3024018,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/21/2024 08:50:18",
      "content": "<p>Given the current API and amount of data, I feel it is very hard to do anything complicated within every batch. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3022073": "Hello everyone,\n\nI'm trying to submit a very light neural network by it timed out. I have no clue of why it didn't work.\n\nI will take any insight if you were successful while submitting with tensorflow or not.\n\n**Related Notebook**: [link](https://www.kaggle.com/code/ulrich07/js-tensorflow-debug/notebook?scriptVersionId=202289829)\n\nHappy Kaggling\n\n**Update**: I found out why it timed out. A single model took ~3 hours to be scored. Since I used 5 fold-models it timed out. But it is very strange that a very light model takes  ~3 hours to be scored.",
    "3022334": "ulrich07 Just to confirm, are you following the rules below?\n\n~~~\nEach batch of predictions (except the very first) must be returned within 10 minutes of the batch features being provided.\n~~~\n\n~~~\nWhen your notebook is run on the hidden test set, inference_server.serve must be called within 15 minutes of the notebook starting or the gateway will throw an error. If you need more than 15 minutes to load your model you can do so during the very first predict call, which does not have the usual 10 minute response deadline.\n~~~",
    "3022446": "Thank you @chumajin. The model is loaded in 5 seconds, and can predict 2M rows in 30 seconds. I will share a notebook.",
    "3022590": "Are you using lags? Also facing a simple Ridge model timing out..",
    "3022669": "Hiya, can you use your GPU allowance? I have definitely verified my account (hence I can take part in this competition) but turning on an accelerator mode doesn't work with the likes of XGBoost with cuda enabled or LGB...",
    "3022776": "I don't think lags is the problem. I have a successfull rige (with and without lags). My TF model do not use lags.",
    "3022893": "try smaller batch size and see how long does it take for your model to process\nyour prediction callback is called on a ~30 batch size for 4mm rows",
    "3023316": "Hi @chumajin  I add the notebook  link, can you have a look at it ? Thks in advance.",
    "3023360": "Did you try removing batch size argument from predict? I ran a simple keras model at one point and didn’t have any issue.",
    "3023361": "Many thanks @simonveitner. I will remove the batch_size parameter.\n\n**Update**: But It's like what you said doesn't work for me. I removed batch_size on a private notebook but I guess it is still not working (submission is still runing for more than 1 hour now). Can you please share a keras code snipet or even check my [notebook](https://www.kaggle.com/code/ulrich07/js-tensorflow-debug).",
    "3023409": "ulrich07 OK. I can see it. I will check it !",
    "3023435": "ulrich07 I created the simulator notebook [here](https://www.kaggle.com/code/chumajin/janestreet-simulator-for-time-series-api). \n\nNow, I'm trying your model, maybe it will take about 2 hours 30 mins.\nSo far, I think your model is fine. Therefore, there might be a bug in the time series API or something error we don't know.\n\nBut I'll run everything through and let you know the results tomorrow morning!",
    "3023462": "ulrich07 could you report here if the model ran through? maybe it's just a time issue.",
    "3023499": "Yes It ran. But didn't check the time probably in ~3 hours",
    "3023580": "ulrich07 I also succeeded with your code(batch size 512). For reference, I changed like this in order to insurance.\n\n~~~\nx = test.select(FE).to_pandas().fillna(3).values\n・・・\npredictions = predictions.fill_nan(0)\n~~~\n\n and the score was -14.4161",
    "3023593": "Thank you @chumajin.  I got -26.xxx score. I am very sad because as it is, putting a TF neural net in the final blend may be problematic concerning timeout issues. But @simonveitner  how long does it take to score your submission. Thks in advance.",
    "3023598": "I think it took less",
    "3023682": "I ran your code and there was no problem. The public score is -16.xxxx and took ~4hrs. \n\nHere is my kernel: https://www.kaggle.com/code/shiyili/js-tensorflow-debug",
    "3024018": "Given the current API and amount of data, I feel it is very hard to do anything complicated within every batch."
  },
  "source": "meta"
}