{
  "id": 546285,
  "title": "Accurate method os submission????",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/546285",
  "author_name": "",
  "post_date": "2024-11-14T20:50:19.669989800Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Can  anybody please tell me how to make correct submission I have trained model and store it .</p>",
  "messages": [
    {
      "id": "3045826",
      "postDate": "11/14/2024 20:50:19",
      "content": "<p>Can  anybody please tell me how to make correct submission I have trained model and store it .</p>",
      "rawMarkdown": "Can  anybody please tell me how to make correct submission I have trained model and store it .",
      "votes": null
    },
    {
      "id": "3045827",
      "postDate": "11/14/2024 20:50:59",
      "content": "<p>I have tried submission other people forked notebook and try to understand it but i have not figured out my mistake.</p>",
      "rawMarkdown": "I have tried submission other people forked notebook and try to understand it but i have not figured out my mistake.",
      "votes": null
    },
    {
      "id": "3046053",
      "postDate": "11/15/2024 04:56:14",
      "content": "<p><a href=\"https://www.kaggle.com/muhammadqasimshabbir\" target=\"_blank\">@muhammadqasimshabbir</a> <br>\nFork the notebook <a href=\"https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission\" target=\"_blank\">here</a> and replace the submission constant with your model. Something like this will work-</p>\n<pre><code>preds = model.predict(yourtestdataset)\npredictions.select(pl.Series(, test[]).with_columns(pl.Series(, preds)).write_parquet()\n</code></pre>",
      "rawMarkdown": "muhammadqasimshabbir \nFork the notebook [here](https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission) and replace the submission constant with your model. Something like this will work-\n\n```python\npreds = model.predict(yourtestdataset)\npredictions.select(pl.Series(\"row_id\", test[\"row_id\"]).with_columns(pl.Series(\"responder_6\", preds)).write_parquet(\"submission.parquet\")\n```",
      "votes": null
    },
    {
      "id": "3046479",
      "postDate": "11/15/2024 14:08:19",
      "content": "<p>thanks i will give it a shot.</p>",
      "rawMarkdown": "thanks i will give it a shot.",
      "votes": null
    },
    {
      "id": "3046481",
      "postDate": "11/15/2024 14:11:33",
      "content": "<p>I have a question here  could you please tell me  your test_data how to is this we have to do the same preprocessing in the <strong><em>def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:</em></strong>  inside this funtion i have to preprocess test data same  as in training  then  I have to input  this dataframe  to model for predcition in above code .  Am right or not ?</p>",
      "rawMarkdown": "I have a question here  could you please tell me  your test_data how to is this we have to do the same preprocessing in the ***def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:***  inside this funtion i have to preprocess test data same  as in training  then  I have to input  this dataframe  to model for predcition in above code .  Am right or not ?",
      "votes": null
    },
    {
      "id": "3046986",
      "postDate": "11/16/2024 05:47:55",
      "content": "<p>Yes of course, this is the same as any other pipeline - you create the preprocessing for the test set batch like the one used in the train set. This is then used for prediction and a new batch is served <a href=\"https://www.kaggle.com/muhammadqasimshabbir\" target=\"_blank\">@muhammadqasimshabbir</a> </p>",
      "rawMarkdown": "Yes of course, this is the same as any other pipeline - you create the preprocessing for the test set batch like the one used in the train set. This is then used for prediction and a new batch is served @muhammadqasimshabbir",
      "votes": null
    },
    {
      "id": "3047404",
      "postDate": "11/16/2024 16:23:23",
      "content": "<p>Thanks for your  help   <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "Thanks for your  help   @ravi20076",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3045827,
      "author_name": "",
      "author_url": "",
      "post_date": "11/14/2024 20:50:59",
      "content": "<p>I have tried submission other people forked notebook and try to understand it but i have not figured out my mistake.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3046053,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "11/15/2024 04:56:14",
          "content": "<p><a href=\"https://www.kaggle.com/muhammadqasimshabbir\" target=\"_blank\">@muhammadqasimshabbir</a> <br>\nFork the notebook <a href=\"https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission\" target=\"_blank\">here</a> and replace the submission constant with your model. Something like this will work-</p>\n<pre><code>preds = model.predict(yourtestdataset)\npredictions.select(pl.Series(, test[]).with_columns(pl.Series(, preds)).write_parquet()\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 3046479,
              "author_name": "",
              "author_url": "",
              "post_date": "11/15/2024 14:08:19",
              "content": "<p>thanks i will give it a shot.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3046481,
                  "author_name": "",
                  "author_url": "",
                  "post_date": "11/15/2024 14:11:33",
                  "content": "<p>I have a question here  could you please tell me  your test_data how to is this we have to do the same preprocessing in the <strong><em>def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:</em></strong>  inside this funtion i have to preprocess test data same  as in training  then  I have to input  this dataframe  to model for predcition in above code .  Am right or not ?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3046986,
                      "author_name": "ravi20076",
                      "author_url": "",
                      "post_date": "11/16/2024 05:47:55",
                      "content": "<p>Yes of course, this is the same as any other pipeline - you create the preprocessing for the test set batch like the one used in the train set. This is then used for prediction and a new batch is served <a href=\"https://www.kaggle.com/muhammadqasimshabbir\" target=\"_blank\">@muhammadqasimshabbir</a> </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3047404,
                          "author_name": "",
                          "author_url": "",
                          "post_date": "11/16/2024 16:23:23",
                          "content": "<p>Thanks for your  help   <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3045826": "Can  anybody please tell me how to make correct submission I have trained model and store it .",
    "3045827": "I have tried submission other people forked notebook and try to understand it but i have not figured out my mistake.",
    "3046053": "muhammadqasimshabbir \nFork the notebook [here](https://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission) and replace the submission constant with your model. Something like this will work-\n\n```python\npreds = model.predict(yourtestdataset)\npredictions.select(pl.Series(\"row_id\", test[\"row_id\"]).with_columns(pl.Series(\"responder_6\", preds)).write_parquet(\"submission.parquet\")\n```",
    "3046479": "thanks i will give it a shot.",
    "3046481": "I have a question here  could you please tell me  your test_data how to is this we have to do the same preprocessing in the ***def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:***  inside this funtion i have to preprocess test data same  as in training  then  I have to input  this dataframe  to model for predcition in above code .  Am right or not ?",
    "3046986": "Yes of course, this is the same as any other pipeline - you create the preprocessing for the test set batch like the one used in the train set. This is then used for prediction and a new batch is served @muhammadqasimshabbir",
    "3047404": "Thanks for your  help   @ravi20076"
  },
  "source": "meta"
}