{
  "id": 543979,
  "title": "Unable to submit!",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543979",
  "author_name": "",
  "post_date": "2024-11-02T16:05:17.421539900Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I first train the predictors of each responder, save them in kaggle/output, and then call them during the prediction process to complete the responder for the test set and predict responder6 based on this. The following is my code：</p>\n<p>`</p>\n<h1>Feature columns setup</h1>\n<p>feature_cols = ['date_id', 'time_id', 'symbol_id', 'weight'] + [f'feature_{i:02}' for i in range(79)]</p>\n<p>def predict(test, lags):<br>\n    # Create a copy of the test dataset for adding predictions of other responders<br>\n    test_copy = test.clone()</p>\n<pre><code> Check if any files are generated in the current directory\n\n\n Use responders 0 to 5 and 7 to 8 for initial predictions\nfor responder in range(9):\n    if responder != 6:   Skip responder_6\n         Load model from /kaggle/working directory dynamically\n        model_path = f\"/kaggle/working/autogluon_responder_model_responder_{responder}\"\n        predictor = TabularPredictor.load(model_path)\n\n        test_copy = test_copy.with_columns(\n            pl.Series(name=f'responder_{responder}', values=predictor.predict(test.to_pandas()[feature_cols]))\n        )\n         Check files again\n        \n\n Use responder_6 for the final prediction\npredictor_6_path = \"/kaggle/working/autogluon_responder_model_responder_6\"\npredictor_6 = TabularPredictor.load(predictor_6_path)\nfeature_cols_with_responders = feature_cols + [f'responder_{i}' for i in range(9) if i != 6]\nresponder_6_predictions = predictor_6.predict(test_copy.to_pandas()[feature_cols_with_responders])\n\ntest = test.with_columns(\n    pl.Series(\n        name=\"responder_6\",\n        values=np.clip(responder_6_predictions, a_min=-5, a_max=5),\n        dtype=pl.Float64\n    )\n)\n\n Combine row_id and responder_6 predictions into the output DataFrame\npredictions = test.select([\"row_id\", \"responder_6\"])\n\n Check if predictions generate any files\ndel test_copy\nprint(\"Files in /kaggle/working after prediction:\", os.listdir(\"/kaggle/working\"))\nprint(predictions.head())\n\n The predict function must return a DataFrame\nassert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n with columns 'row_id', 'responder_6'\nassert predictions.columns == ['row_id', 'responder_6']\n and as many rows as the test data.\nassert len(predictions) == len(test)\n\nreturn predictions\n</code></pre>\n<p>inference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)</p>\n<p>if os.getenv('KAGGLE_IS_COMPETITION_RERUN'):<br>\n    inference_server.serve()<br>\nelse:<br>\n    inference_server.run_local_gateway(<br>\n        (<br>\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',<br>\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',<br>\n        )<br>\n    )<br>\n`</p>",
  "messages": [
    {
      "id": "3034766",
      "postDate": "11/02/2024 16:05:17",
      "content": "<p>I first train the predictors of each responder, save them in kaggle/output, and then call them during the prediction process to complete the responder for the test set and predict responder6 based on this. The following is my code：</p>\n<p>`</p>\n<h1>Feature columns setup</h1>\n<p>feature_cols = ['date_id', 'time_id', 'symbol_id', 'weight'] + [f'feature_{i:02}' for i in range(79)]</p>\n<p>def predict(test, lags):<br>\n    # Create a copy of the test dataset for adding predictions of other responders<br>\n    test_copy = test.clone()</p>\n<pre><code> Check if any files are generated in the current directory\n\n\n Use responders 0 to 5 and 7 to 8 for initial predictions\nfor responder in range(9):\n    if responder != 6:   Skip responder_6\n         Load model from /kaggle/working directory dynamically\n        model_path = f\"/kaggle/working/autogluon_responder_model_responder_{responder}\"\n        predictor = TabularPredictor.load(model_path)\n\n        test_copy = test_copy.with_columns(\n            pl.Series(name=f'responder_{responder}', values=predictor.predict(test.to_pandas()[feature_cols]))\n        )\n         Check files again\n        \n\n Use responder_6 for the final prediction\npredictor_6_path = \"/kaggle/working/autogluon_responder_model_responder_6\"\npredictor_6 = TabularPredictor.load(predictor_6_path)\nfeature_cols_with_responders = feature_cols + [f'responder_{i}' for i in range(9) if i != 6]\nresponder_6_predictions = predictor_6.predict(test_copy.to_pandas()[feature_cols_with_responders])\n\ntest = test.with_columns(\n    pl.Series(\n        name=\"responder_6\",\n        values=np.clip(responder_6_predictions, a_min=-5, a_max=5),\n        dtype=pl.Float64\n    )\n)\n\n Combine row_id and responder_6 predictions into the output DataFrame\npredictions = test.select([\"row_id\", \"responder_6\"])\n\n Check if predictions generate any files\ndel test_copy\nprint(\"Files in /kaggle/working after prediction:\", os.listdir(\"/kaggle/working\"))\nprint(predictions.head())\n\n The predict function must return a DataFrame\nassert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n with columns 'row_id', 'responder_6'\nassert predictions.columns == ['row_id', 'responder_6']\n and as many rows as the test data.\nassert len(predictions) == len(test)\n\nreturn predictions\n</code></pre>\n<p>inference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)</p>\n<p>if os.getenv('KAGGLE_IS_COMPETITION_RERUN'):<br>\n    inference_server.serve()<br>\nelse:<br>\n    inference_server.run_local_gateway(<br>\n        (<br>\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',<br>\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',<br>\n        )<br>\n    )<br>\n`</p>",
      "rawMarkdown": "I first train the predictors of each responder, save them in kaggle/output, and then call them during the prediction process to complete the responder for the test set and predict responder6 based on this. The following is my code：\n\n`\n# Feature columns setup\nfeature_cols = ['date_id', 'time_id', 'symbol_id', 'weight'] + [f'feature_{i:02}' for i in range(79)]\n\ndef predict(test, lags):\n    # Create a copy of the test dataset for adding predictions of other responders\n    test_copy = test.clone()\n    \n    # Check if any files are generated in the current directory\n    #print(\"Files in /kaggle/working:\", os.listdir(\"/kaggle/working\"))\n    \n    # Use responders 0 to 5 and 7 to 8 for initial predictions\n    for responder in range(9):\n        if responder != 6:  # Skip responder_6\n            # Load model from /kaggle/working directory dynamically\n            model_path = f\"/kaggle/working/autogluon_responder_model_responder_{responder}\"\n            predictor = TabularPredictor.load(model_path)\n            \n            test_copy = test_copy.with_columns(\n                pl.Series(name=f'responder_{responder}', values=predictor.predict(test.to_pandas()[feature_cols]))\n            )\n            # Check files again\n            #print(f\"After predictor {responder}, files in /kaggle/working:\", os.listdir(\"/kaggle/working\"))\n\n    # Use responder_6 for the final prediction\n    predictor_6_path = \"/kaggle/working/autogluon_responder_model_responder_6\"\n    predictor_6 = TabularPredictor.load(predictor_6_path)\n    feature_cols_with_responders = feature_cols + [f'responder_{i}' for i in range(9) if i != 6]\n    responder_6_predictions = predictor_6.predict(test_copy.to_pandas()[feature_cols_with_responders])\n\n    test = test.with_columns(\n        pl.Series(\n            name=\"responder_6\",\n            values=np.clip(responder_6_predictions, a_min=-5, a_max=5),\n            dtype=pl.Float64\n        )\n    )\n\n    # Combine row_id and responder_6 predictions into the output DataFrame\n    predictions = test.select([\"row_id\", \"responder_6\"])\n\n    # Check if predictions generate any files\n    del test_copy\n    print(\"Files in /kaggle/working after prediction:\", os.listdir(\"/kaggle/working\"))\n    print(predictions.head())\n    \n    # The predict function must return a DataFrame\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n    # with columns 'row_id', 'responder_6'\n    assert predictions.columns == ['row_id', 'responder_6']\n    # and as many rows as the test data.\n    assert len(predictions) == len(test)\n\n    return predictions\n\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n        )\n    )\n`",
      "votes": null
    },
    {
      "id": "3034768",
      "postDate": "11/02/2024 16:07:47",
      "content": "<p>This is my first official participation in the competition. I may ask some very novice questions (such as this). If anyone can help, I will be very grateful! ! !</p>",
      "rawMarkdown": "This is my first official participation in the competition. I may ask some very novice questions (such as this). If anyone can help, I will be very grateful! ! !",
      "votes": null
    },
    {
      "id": "3038524",
      "postDate": "11/07/2024 04:55:30",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/dearluna\" target=\"_blank\">@dearluna</a> </p>\n<p>Did you manage to make a submission?</p>",
      "rawMarkdown": "Hey @dearluna \n\nDid you manage to make a submission?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3034768,
      "author_name": "dearluna",
      "author_url": "",
      "post_date": "11/02/2024 16:07:47",
      "content": "<p>This is my first official participation in the competition. I may ask some very novice questions (such as this). If anyone can help, I will be very grateful! ! !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3038524,
      "author_name": "kaleab7",
      "author_url": "",
      "post_date": "11/07/2024 04:55:30",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/dearluna\" target=\"_blank\">@dearluna</a> </p>\n<p>Did you manage to make a submission?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3034766": "I first train the predictors of each responder, save them in kaggle/output, and then call them during the prediction process to complete the responder for the test set and predict responder6 based on this. The following is my code：\n\n`\n# Feature columns setup\nfeature_cols = ['date_id', 'time_id', 'symbol_id', 'weight'] + [f'feature_{i:02}' for i in range(79)]\n\ndef predict(test, lags):\n    # Create a copy of the test dataset for adding predictions of other responders\n    test_copy = test.clone()\n    \n    # Check if any files are generated in the current directory\n    #print(\"Files in /kaggle/working:\", os.listdir(\"/kaggle/working\"))\n    \n    # Use responders 0 to 5 and 7 to 8 for initial predictions\n    for responder in range(9):\n        if responder != 6:  # Skip responder_6\n            # Load model from /kaggle/working directory dynamically\n            model_path = f\"/kaggle/working/autogluon_responder_model_responder_{responder}\"\n            predictor = TabularPredictor.load(model_path)\n            \n            test_copy = test_copy.with_columns(\n                pl.Series(name=f'responder_{responder}', values=predictor.predict(test.to_pandas()[feature_cols]))\n            )\n            # Check files again\n            #print(f\"After predictor {responder}, files in /kaggle/working:\", os.listdir(\"/kaggle/working\"))\n\n    # Use responder_6 for the final prediction\n    predictor_6_path = \"/kaggle/working/autogluon_responder_model_responder_6\"\n    predictor_6 = TabularPredictor.load(predictor_6_path)\n    feature_cols_with_responders = feature_cols + [f'responder_{i}' for i in range(9) if i != 6]\n    responder_6_predictions = predictor_6.predict(test_copy.to_pandas()[feature_cols_with_responders])\n\n    test = test.with_columns(\n        pl.Series(\n            name=\"responder_6\",\n            values=np.clip(responder_6_predictions, a_min=-5, a_max=5),\n            dtype=pl.Float64\n        )\n    )\n\n    # Combine row_id and responder_6 predictions into the output DataFrame\n    predictions = test.select([\"row_id\", \"responder_6\"])\n\n    # Check if predictions generate any files\n    del test_copy\n    print(\"Files in /kaggle/working after prediction:\", os.listdir(\"/kaggle/working\"))\n    print(predictions.head())\n    \n    # The predict function must return a DataFrame\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n    # with columns 'row_id', 'responder_6'\n    assert predictions.columns == ['row_id', 'responder_6']\n    # and as many rows as the test data.\n    assert len(predictions) == len(test)\n\n    return predictions\n\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n        )\n    )\n`",
    "3034768": "This is my first official participation in the competition. I may ask some very novice questions (such as this). If anyone can help, I will be very grateful! ! !",
    "3038524": "Hey @dearluna \n\nDid you manage to make a submission?"
  },
  "source": "meta"
}