{
  "id": 542363,
  "title": "How to avoid exceeding the time limit?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/542363",
  "author_name": "",
  "post_date": "2024-10-24T12:29:07.570061100Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I have currently used 3 million data and an LGB model with 20 iterators and 2 folds. I completed training and inference offline in 100 seconds, but due to the 4.5 million test data and batch loading, it took me 4 hours to get score. Currently, it has only reached 0.0025, and I don't know how to improve the performance.</p>\n<p>The offline training took over 100 seconds, indicating that training inference separation should be useless.</p>",
  "messages": [
    {
      "id": "3026999",
      "postDate": "10/24/2024 12:29:07",
      "content": "<p>I have currently used 3 million data and an LGB model with 20 iterators and 2 folds. I completed training and inference offline in 100 seconds, but due to the 4.5 million test data and batch loading, it took me 4 hours to get score. Currently, it has only reached 0.0025, and I don't know how to improve the performance.</p>\n<p>The offline training took over 100 seconds, indicating that training inference separation should be useless.</p>",
      "rawMarkdown": "I have currently used 3 million data and an LGB model with 20 iterators and 2 folds. I completed training and inference offline in 100 seconds, but due to the 4.5 million test data and batch loading, it took me 4 hours to get score. Currently, it has only reached 0.0025, and I don't know how to improve the performance.\n\nThe offline training took over 100 seconds, indicating that training inference separation should be useless.",
      "votes": null
    },
    {
      "id": "3027004",
      "postDate": "10/24/2024 12:34:30",
      "content": "<p>You should try to optimize your code for fast inference. Use the offline api to analyze which part takes time during inference.</p>",
      "rawMarkdown": "You should try to optimize your code for fast inference. Use the offline api to analyze which part takes time during inference.",
      "votes": null
    },
    {
      "id": "3027005",
      "postDate": "10/24/2024 12:34:33",
      "content": "<p>Did you do very heavy feature engineering during the inference? In my case, a simple one-fold LGB model (using the native 79 features) takes around 30min to score in the LB. That is a lot difference to 4hrs. </p>",
      "rawMarkdown": "Did you do very heavy feature engineering during the inference? In my case, a simple one-fold LGB model (using the native 79 features) takes around 30min to score in the LB. That is a lot difference to 4hrs.",
      "votes": null
    },
    {
      "id": "3027020",
      "postDate": "10/24/2024 12:49:23",
      "content": "<p>Excuse me, have you turned on the GPU? How much data did you use?</p>",
      "rawMarkdown": "Excuse me, have you turned on the GPU? How much data did you use?",
      "votes": null
    },
    {
      "id": "3027033",
      "postDate": "10/24/2024 13:13:51",
      "content": "<p>No GPU. My code has only inference. Training used all data in partition 6-9 except the las 100 days as validation.</p>\n<p>You can use my synthetic test data to estimate your submission scoring time. The full score time should be around 30x longer than the time of inference using the synthetic test data.</p>",
      "rawMarkdown": "No GPU. My code has only inference. Training used all data in partition 6-9 except the las 100 days as validation.\n\nYou can use my synthetic test data to estimate your submission scoring time. The full score time should be around 30x longer than the time of inference using the synthetic test data.",
      "votes": null
    },
    {
      "id": "3027045",
      "postDate": "10/24/2024 13:31:16",
      "content": "<p>thank you,I will have a try.</p>",
      "rawMarkdown": "thank you,I will have a try.",
      "votes": null
    },
    {
      "id": "3027643",
      "postDate": "10/25/2024 04:32:45",
      "content": "<p>How can we do that?? My notebook had just two things. Simple imputer and prediction, which took 17.8 seconds to run and save the notebook. However, it took about 1hr to get the scores.</p>",
      "rawMarkdown": "How can we do that?? My notebook had just two things. Simple imputer and prediction, which took 17.8 seconds to run and save the notebook. However, it took about 1hr to get the scores.",
      "votes": null
    },
    {
      "id": "3027653",
      "postDate": "10/25/2024 04:56:28",
      "content": "<p>Less number of iterations and use a more optimized model for running the inference </p>",
      "rawMarkdown": "Less number of iterations and use a more optimized model for running the inference",
      "votes": null
    },
    {
      "id": "3027775",
      "postDate": "10/25/2024 09:10:22",
      "content": "<p>Has anyone tested how long the API takes to evaluate a constant prediction of 0 and whether that time is consistent or fluctuates significantly? This time serves as the baseline, representing the API's inherent loading time. I am particularly concerned that the duration may depend on the global load, which would be uncontrollable.</p>",
      "rawMarkdown": "Has anyone tested how long the API takes to evaluate a constant prediction of 0 and whether that time is consistent or fluctuates significantly? This time serves as the baseline, representing the API's inherent loading time. I am particularly concerned that the duration may depend on the global load, which would be uncontrollable.",
      "votes": null
    },
    {
      "id": "3028543",
      "postDate": "10/26/2024 08:26:24",
      "content": "<p>roughly 9 minutes, and yes it fluctuates (sometimes 10). </p>",
      "rawMarkdown": "roughly 9 minutes, and yes it fluctuates (sometimes 10).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3027004,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "10/24/2024 12:34:30",
      "content": "<p>You should try to optimize your code for fast inference. Use the offline api to analyze which part takes time during inference.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3027643,
          "author_name": "jayshrivastava",
          "author_url": "",
          "post_date": "10/25/2024 04:32:45",
          "content": "<p>How can we do that?? My notebook had just two things. Simple imputer and prediction, which took 17.8 seconds to run and save the notebook. However, it took about 1hr to get the scores.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3027005,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/24/2024 12:34:33",
      "content": "<p>Did you do very heavy feature engineering during the inference? In my case, a simple one-fold LGB model (using the native 79 features) takes around 30min to score in the LB. That is a lot difference to 4hrs. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3027020,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "10/24/2024 12:49:23",
          "content": "<p>Excuse me, have you turned on the GPU? How much data did you use?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3027033,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "10/24/2024 13:13:51",
              "content": "<p>No GPU. My code has only inference. Training used all data in partition 6-9 except the las 100 days as validation.</p>\n<p>You can use my synthetic test data to estimate your submission scoring time. The full score time should be around 30x longer than the time of inference using the synthetic test data.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3027045,
                  "author_name": "yunsuxiaozi",
                  "author_url": "",
                  "post_date": "10/24/2024 13:31:16",
                  "content": "<p>thank you,I will have a try.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3027653,
      "author_name": "skylord",
      "author_url": "",
      "post_date": "10/25/2024 04:56:28",
      "content": "<p>Less number of iterations and use a more optimized model for running the inference </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3027775,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "10/25/2024 09:10:22",
      "content": "<p>Has anyone tested how long the API takes to evaluate a constant prediction of 0 and whether that time is consistent or fluctuates significantly? This time serves as the baseline, representing the API's inherent loading time. I am particularly concerned that the duration may depend on the global load, which would be uncontrollable.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3028543,
          "author_name": "lucasmorin",
          "author_url": "",
          "post_date": "10/26/2024 08:26:24",
          "content": "<p>roughly 9 minutes, and yes it fluctuates (sometimes 10). </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3026999": "I have currently used 3 million data and an LGB model with 20 iterators and 2 folds. I completed training and inference offline in 100 seconds, but due to the 4.5 million test data and batch loading, it took me 4 hours to get score. Currently, it has only reached 0.0025, and I don't know how to improve the performance.\n\nThe offline training took over 100 seconds, indicating that training inference separation should be useless.",
    "3027004": "You should try to optimize your code for fast inference. Use the offline api to analyze which part takes time during inference.",
    "3027005": "Did you do very heavy feature engineering during the inference? In my case, a simple one-fold LGB model (using the native 79 features) takes around 30min to score in the LB. That is a lot difference to 4hrs.",
    "3027020": "Excuse me, have you turned on the GPU? How much data did you use?",
    "3027033": "No GPU. My code has only inference. Training used all data in partition 6-9 except the las 100 days as validation.\n\nYou can use my synthetic test data to estimate your submission scoring time. The full score time should be around 30x longer than the time of inference using the synthetic test data.",
    "3027045": "thank you,I will have a try.",
    "3027643": "How can we do that?? My notebook had just two things. Simple imputer and prediction, which took 17.8 seconds to run and save the notebook. However, it took about 1hr to get the scores.",
    "3027653": "Less number of iterations and use a more optimized model for running the inference",
    "3027775": "Has anyone tested how long the API takes to evaluate a constant prediction of 0 and whether that time is consistent or fluctuates significantly? This time serves as the baseline, representing the API's inherent loading time. I am particularly concerned that the duration may depend on the global load, which would be uncontrollable.",
    "3028543": "roughly 9 minutes, and yes it fluctuates (sometimes 10)."
  },
  "source": "meta"
}