{
  "id": 547416,
  "title": "Trouble with scoring",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/547416",
  "author_name": "",
  "post_date": "2024-11-21T12:06:47.681721300Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I did a CatBoost model which runs in a couple of minutes and gives a reasonable 'submission.parquet' file, so no trouble with running the notebook, but all five attempts failed to score. Can you please advise me? <br>\nThis is my predict function:</p>\n<p>lags_ : pl.DataFrame | None = None<br>\nimport numpy as np</p>\n<h1>Replace this function with your inference code.</h1>\n<h1>You can return either a Pandas or Polars dataframe, though Polars is recommended.</h1>\n<h1>Each batch of predictions (except the very first) must be returned within 1 minute of the batch features being provided.</h1>\n<p>def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:<br>\n    \"\"\"Make a prediction.\"\"\"<br>\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.<br>\n    # Use them as extra features, if you like.<br>\n    global lags_<br>\n    if lags is not None:<br>\n        lags_ = lags<br>\n    lags = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet/date_id=0')<br>\n    test = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0')<br>\n    test['date_id'] = test['date_id'].astype('int16')</p>\n<pre><code>test[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest = test.drop(columns = [, ])\n\n\ncolumns_to_select = list(range(,)) + [, ]\n\ntest.iloc[:, :] = ().fit_transform(test.iloc[:, :])\ntest.iloc[:,:] = ().fit_transform(test.iloc[:,:])\ndata_matrix = test.values\n\n# : feature matrix (for this analysis, we exclude the  variable from the dataset)\nx_test = data_matrix[:, columns_to_select]\n\n# y: labels vector\ny_test = data_matrix[:, ]\ny_pred = cb_cls.predict(x_test)\ny_pred = np.round(y_pred, decimals=)\ny_pred = y_pred.tolist()\n#  this section with your own predictions\npredictions = pd.({: test[], : y_pred}) \n\npredictions.to_parquet(, index=)\n\nif isinstance(predictions, pl.):\n    assert predictions.columns == [, ]\nelif isinstance(predictions, pd.):\n    assert (predictions.columns == [, ]).all()\nelse:\n    raise ()\n#  has as many rows as the test data.\nassert len(predictions) == len(test)\n\nreturn predictions\n</code></pre>",
  "messages": [
    {
      "id": "3051580",
      "postDate": "11/21/2024 12:06:47",
      "content": "<p>Hi,</p>\n<p>I did a CatBoost model which runs in a couple of minutes and gives a reasonable 'submission.parquet' file, so no trouble with running the notebook, but all five attempts failed to score. Can you please advise me? <br>\nThis is my predict function:</p>\n<p>lags_ : pl.DataFrame | None = None<br>\nimport numpy as np</p>\n<h1>Replace this function with your inference code.</h1>\n<h1>You can return either a Pandas or Polars dataframe, though Polars is recommended.</h1>\n<h1>Each batch of predictions (except the very first) must be returned within 1 minute of the batch features being provided.</h1>\n<p>def predict(test: pl.DataFrame, lags: pl.DataFrame | None) -&gt; pl.DataFrame | pd.DataFrame:<br>\n    \"\"\"Make a prediction.\"\"\"<br>\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.<br>\n    # Use them as extra features, if you like.<br>\n    global lags_<br>\n    if lags is not None:<br>\n        lags_ = lags<br>\n    lags = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet/date_id=0')<br>\n    test = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0')<br>\n    test['date_id'] = test['date_id'].astype('int16')</p>\n<pre><code>test[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest[] = lags[]\ntest = test.drop(columns = [, ])\n\n\ncolumns_to_select = list(range(,)) + [, ]\n\ntest.iloc[:, :] = ().fit_transform(test.iloc[:, :])\ntest.iloc[:,:] = ().fit_transform(test.iloc[:,:])\ndata_matrix = test.values\n\n# : feature matrix (for this analysis, we exclude the  variable from the dataset)\nx_test = data_matrix[:, columns_to_select]\n\n# y: labels vector\ny_test = data_matrix[:, ]\ny_pred = cb_cls.predict(x_test)\ny_pred = np.round(y_pred, decimals=)\ny_pred = y_pred.tolist()\n#  this section with your own predictions\npredictions = pd.({: test[], : y_pred}) \n\npredictions.to_parquet(, index=)\n\nif isinstance(predictions, pl.):\n    assert predictions.columns == [, ]\nelif isinstance(predictions, pd.):\n    assert (predictions.columns == [, ]).all()\nelse:\n    raise ()\n#  has as many rows as the test data.\nassert len(predictions) == len(test)\n\nreturn predictions\n</code></pre>",
      "rawMarkdown": "Hi,\n\nI did a CatBoost model which runs in a couple of minutes and gives a reasonable 'submission.parquet' file, so no trouble with running the notebook, but all five attempts failed to score. Can you please advise me? \nThis is my predict function:\n\nlags_ : pl.DataFrame | None = None\nimport numpy as np\n\n# Replace this function with your inference code.\n# You can return either a Pandas or Polars dataframe, though Polars is recommended.\n# Each batch of predictions (except the very first) must be returned within 1 minute of the batch features being provided.\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction.\"\"\"\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n    # Use them as extra features, if you like.\n    global lags_\n    if lags is not None:\n        lags_ = lags\n    lags = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet/date_id=0')\n    test = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0')\n    test['date_id'] = test['date_id'].astype('int16')\n\n\n    test['responder_0'] = lags['responder_0_lag_1']\n    test['responder_1'] = lags['responder_1_lag_1']\n    test['responder_2'] = lags['responder_2_lag_1']\n    test['responder_3'] = lags['responder_3_lag_1']\n    test['responder_4'] = lags['responder_4_lag_1']\n    test['responder_5'] = lags['responder_5_lag_1']\n    test['responder_6'] = lags['responder_6_lag_1']\n    test['responder_7'] = lags['responder_7_lag_1']\n    test['responder_8'] = lags['responder_8_lag_1']\n    test = test.drop(columns = ['weight', 'is_scored'])\n\n\n    columns_to_select = list(range(1,89)) + [90, 91]\n\n    test.iloc[:, 4:89] = StandardScaler().fit_transform(test.iloc[:, 4:89])\n    test.iloc[:,90:] = StandardScaler().fit_transform(test.iloc[:,90:])\n    data_matrix = test.values\n\n    # X: feature matrix (for this analysis, we exclude the 'revenue' variable from the dataset)\n    x_test = data_matrix[:, columns_to_select]\n\n    # y: labels vector\n    y_test = data_matrix[:, 89]\n    y_pred = cb_cls.predict(x_test)\n    y_pred = np.round(y_pred, decimals=1)\n    y_pred = y_pred.tolist()\n    # Replace this section with your own predictions\n    predictions = pd.DataFrame({'row_id': test['row_id'], 'responder_6': y_pred}) \n   \n    predictions.to_parquet('submission.parquet', index=False)\n  \n    if isinstance(predictions, pl.DataFrame):\n        assert predictions.columns == ['row_id', 'responder_6']\n    elif isinstance(predictions, pd.DataFrame):\n        assert (predictions.columns == ['row_id', 'responder_6']).all()\n    else:\n        raise TypeError('The predict function must return a DataFrame to me')\n    # Confirm has as many rows as the test data.\n    assert len(predictions) == len(test)\n\n    return predictions",
      "votes": null
    },
    {
      "id": "3051584",
      "postDate": "11/21/2024 12:11:44",
      "content": "<p>Are you trying to load the test data in your predict function?</p>\n<pre><code>lags = pd.read_parquet()\ntest = pd.read_parquet()\n</code></pre>\n<p>These lines here should not be there. Even for the test data the server will pass you the data via the test parameter to your predict function.</p>\n<p>You should use the dataframes passed in as arguments to predict.</p>",
      "rawMarkdown": "Are you trying to load the test data in your predict function?\n\n\n```python\nlags = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet/date_id=0')\ntest = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0')\n```\nThese lines here should not be there. Even for the test data the server will pass you the data via the test parameter to your predict function.\n\nYou should use the dataframes passed in as arguments to predict.",
      "votes": null
    },
    {
      "id": "3051778",
      "postDate": "11/21/2024 16:14:01",
      "content": "<p>Michael,</p>\n<p>Thank you, your comment actually worked! I replaced this with <br>\ntest_df = test.to_pandas()<br>\nlags_df = lags_.to_pandas()<br>\n(I looked at discussion post by John Starfield and found those lines there)<br>\nand everything started to score!</p>\n<p>Thank you again!</p>",
      "rawMarkdown": "Michael,\n\nThank you, your comment actually worked! I replaced this with \ntest_df = test.to_pandas()\nlags_df = lags_.to_pandas()\n(I looked at discussion post by John Starfield and found those lines there)\nand everything started to score!\n\nThank you again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3051584,
      "author_name": "michaeltimbs",
      "author_url": "",
      "post_date": "11/21/2024 12:11:44",
      "content": "<p>Are you trying to load the test data in your predict function?</p>\n<pre><code>lags = pd.read_parquet()\ntest = pd.read_parquet()\n</code></pre>\n<p>These lines here should not be there. Even for the test data the server will pass you the data via the test parameter to your predict function.</p>\n<p>You should use the dataframes passed in as arguments to predict.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3051778,
      "author_name": "anastasiaruzmaikina",
      "author_url": "",
      "post_date": "11/21/2024 16:14:01",
      "content": "<p>Michael,</p>\n<p>Thank you, your comment actually worked! I replaced this with <br>\ntest_df = test.to_pandas()<br>\nlags_df = lags_.to_pandas()<br>\n(I looked at discussion post by John Starfield and found those lines there)<br>\nand everything started to score!</p>\n<p>Thank you again!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3051580": "Hi,\n\nI did a CatBoost model which runs in a couple of minutes and gives a reasonable 'submission.parquet' file, so no trouble with running the notebook, but all five attempts failed to score. Can you please advise me? \nThis is my predict function:\n\nlags_ : pl.DataFrame | None = None\nimport numpy as np\n\n# Replace this function with your inference code.\n# You can return either a Pandas or Polars dataframe, though Polars is recommended.\n# Each batch of predictions (except the very first) must be returned within 1 minute of the batch features being provided.\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction.\"\"\"\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n    # Use them as extra features, if you like.\n    global lags_\n    if lags is not None:\n        lags_ = lags\n    lags = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet/date_id=0')\n    test = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0')\n    test['date_id'] = test['date_id'].astype('int16')\n\n\n    test['responder_0'] = lags['responder_0_lag_1']\n    test['responder_1'] = lags['responder_1_lag_1']\n    test['responder_2'] = lags['responder_2_lag_1']\n    test['responder_3'] = lags['responder_3_lag_1']\n    test['responder_4'] = lags['responder_4_lag_1']\n    test['responder_5'] = lags['responder_5_lag_1']\n    test['responder_6'] = lags['responder_6_lag_1']\n    test['responder_7'] = lags['responder_7_lag_1']\n    test['responder_8'] = lags['responder_8_lag_1']\n    test = test.drop(columns = ['weight', 'is_scored'])\n\n\n    columns_to_select = list(range(1,89)) + [90, 91]\n\n    test.iloc[:, 4:89] = StandardScaler().fit_transform(test.iloc[:, 4:89])\n    test.iloc[:,90:] = StandardScaler().fit_transform(test.iloc[:,90:])\n    data_matrix = test.values\n\n    # X: feature matrix (for this analysis, we exclude the 'revenue' variable from the dataset)\n    x_test = data_matrix[:, columns_to_select]\n\n    # y: labels vector\n    y_test = data_matrix[:, 89]\n    y_pred = cb_cls.predict(x_test)\n    y_pred = np.round(y_pred, decimals=1)\n    y_pred = y_pred.tolist()\n    # Replace this section with your own predictions\n    predictions = pd.DataFrame({'row_id': test['row_id'], 'responder_6': y_pred}) \n   \n    predictions.to_parquet('submission.parquet', index=False)\n  \n    if isinstance(predictions, pl.DataFrame):\n        assert predictions.columns == ['row_id', 'responder_6']\n    elif isinstance(predictions, pd.DataFrame):\n        assert (predictions.columns == ['row_id', 'responder_6']).all()\n    else:\n        raise TypeError('The predict function must return a DataFrame to me')\n    # Confirm has as many rows as the test data.\n    assert len(predictions) == len(test)\n\n    return predictions",
    "3051584": "Are you trying to load the test data in your predict function?\n\n\n```python\nlags = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet/date_id=0')\ntest = pd.read_parquet('/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0')\n```\nThese lines here should not be there. Even for the test data the server will pass you the data via the test parameter to your predict function.\n\nYou should use the dataframes passed in as arguments to predict.",
    "3051778": "Michael,\n\nThank you, your comment actually worked! I replaced this with \ntest_df = test.to_pandas()\nlags_df = lags_.to_pandas()\n(I looked at discussion post by John Starfield and found those lines there)\nand everything started to score!\n\nThank you again!"
  },
  "source": "meta"
}