{
  "id": 548586,
  "title": "Why always \"Submission Format Error\"",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/548586",
  "author_name": "Mr Lin",
  "post_date": "2024-11-27T15:35:27.240000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Could please help me identify what's wrong with my code? I've tried over 30 times but continue to encounter a \"Submission Format Error.\"</p>\n<pre><code>model = create_model()\n\nmodel_paths = [\n    ,\n    ,\n    ,\n    ,\n    \n]\n</code></pre>\n<pre><code> sklearn.preprocessing  StandardScaler\n\nlags_ : pl.DataFrame |  = \n\n () -&gt; pl.DataFrame | pd.DataFrame:\n    \n    \n    \n     lags_\n     lags   :\n        lags_ = lags\n\n    test = test.fill_null()\n    test = test.join(lags_, on=[, , ], how=)\n\n    predictions = test.select(\n        ,\n        pl.lit().alias()\n    )\n\n    feature_input = test.select(\n    [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ]\n    )\n\n    preds = []\n     model_path  model_paths:\n        model.load_weights(model_path)  \n        prediction = model.predict(feature_input)  \n        preds.append(prediction)\n\n    final_pred = (preds) / (preds)\n\n\n    final_pred_df = pd.DataFrame(final_pred, columns=[])\n    test_df = pd.DataFrame(test[], columns=[])\n\n    predictions = pd.DataFrame({\n    : test_df[].values,  \n    : final_pred_df[].values  \n    })\n\n\n    \n     (predictions, pl.DataFrame | pd.DataFrame)\n\n    \n     (predictions.columns) == [, ]\n\n    \n     (predictions) == (test)\n\n     predictions\n</code></pre>",
  "messages": [
    {
      "id": 3056980,
      "postDate": "2024-11-27T15:35:27.240Z",
      "content": "<p>Could please help me identify what's wrong with my code? I've tried over 30 times but continue to encounter a \"Submission Format Error.\"</p>\n<pre><code>model = create_model()\n\nmodel_paths = [\n    ,\n    ,\n    ,\n    ,\n    \n]\n</code></pre>\n<pre><code> sklearn.preprocessing  StandardScaler\n\nlags_ : pl.DataFrame |  = \n\n () -&gt; pl.DataFrame | pd.DataFrame:\n    \n    \n    \n     lags_\n     lags   :\n        lags_ = lags\n\n    test = test.fill_null()\n    test = test.join(lags_, on=[, , ], how=)\n\n    predictions = test.select(\n        ,\n        pl.lit().alias()\n    )\n\n    feature_input = test.select(\n    [, , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ]\n    )\n\n    preds = []\n     model_path  model_paths:\n        model.load_weights(model_path)  \n        prediction = model.predict(feature_input)  \n        preds.append(prediction)\n\n    final_pred = (preds) / (preds)\n\n\n    final_pred_df = pd.DataFrame(final_pred, columns=[])\n    test_df = pd.DataFrame(test[], columns=[])\n\n    predictions = pd.DataFrame({\n    : test_df[].values,  \n    : final_pred_df[].values  \n    })\n\n\n    \n     (predictions, pl.DataFrame | pd.DataFrame)\n\n    \n     (predictions.columns) == [, ]\n\n    \n     (predictions) == (test)\n\n     predictions\n</code></pre>",
      "rawMarkdown": "Could please help me identify what's wrong with my code? I've tried over 30 times but continue to encounter a \"Submission Format Error.\"\n\n```python\nmodel = create_model()\n\nmodel_paths = [\n    \"/kaggle/input/jane-street/other/default/5/model_0.weights.h5\",\n    \"/kaggle/input/jane-street/other/default/5/model_1.weights.h5\",\n    \"/kaggle/input/jane-street/other/default/5/model_2.weights.h5\",\n    \"/kaggle/input/jane-street/other/default/5/model_3.weights.h5\",\n    \"/kaggle/input/jane-street/other/default/5/model_4.weights.h5\"\n]\n```\n```python\nfrom sklearn.preprocessing import StandardScaler\n\nlags_ : pl.DataFrame | None = None\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction.\"\"\"\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n    # Use them as extra features, if you like.\n    global lags_\n    if lags is not None:\n        lags_ = lags\n    \n    test = test.fill_null(0)\n    test = test.join(lags_, on=[\"date_id\", \"time_id\", \"symbol_id\"], how=\"inner\")\n\n    predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6')\n    )\n\n    feature_input = test.select(\n    ['symbol_id', 'weight', 'feature_00', 'feature_01', 'feature_02', 'feature_03', 'feature_04', 'feature_05', 'feature_06', 'feature_07', 'feature_08', 'feature_09', 'feature_10', 'feature_11', 'feature_12', 'feature_13', 'feature_14', 'feature_15', 'feature_16', 'feature_17', 'feature_18', 'feature_19', 'feature_20', 'feature_21', 'feature_22', 'feature_23', 'feature_24', 'feature_25', 'feature_26', 'feature_27', 'feature_28', 'feature_29', 'feature_30', 'feature_31', 'feature_32', 'feature_33', 'feature_34', 'feature_35', 'feature_36', 'feature_37', 'feature_38', 'feature_39', 'feature_40', 'feature_41', 'feature_42', 'feature_43', 'feature_44', 'feature_45', 'feature_46', 'feature_47', 'feature_48', 'feature_49', 'feature_50', 'feature_51', 'feature_52', 'feature_53', 'feature_54', 'feature_55', 'feature_56', 'feature_57', 'feature_58', 'feature_59', 'feature_60', 'feature_61', 'feature_62', 'feature_63', 'feature_64', 'feature_65', 'feature_66', 'feature_67', 'feature_68', 'feature_69', 'feature_70', 'feature_71', 'feature_72', 'feature_73', 'feature_74', 'feature_75', 'feature_76', 'feature_77', 'feature_78', 'responder_0_lag_1', 'responder_1_lag_1', 'responder_2_lag_1', 'responder_3_lag_1', 'responder_4_lag_1', 'responder_5_lag_1']\n    )\n\n    preds = []\n    for model_path in model_paths:\n        model.load_weights(model_path)  \n        prediction = model.predict(feature_input)  \n        preds.append(prediction)\n    \n    final_pred = sum(preds) / len(preds)\n    \n\n    final_pred_df = pd.DataFrame(final_pred, columns=[\"predictions\"])\n    test_df = pd.DataFrame(test[\"row_id\"], columns=[\"row_id\"])\n    \n    predictions = pd.DataFrame({\n    \"row_id\": test_df[\"row_id\"].values,  \n    \"responder_6\": final_pred_df[\"predictions\"].values  \n    })\n    \n  \n    # Ensure the prediction function returns a DataFrame\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n\n    # Ensure the returned DataFrame has columns 'row_id' and 'responder_6'\n    assert list(predictions.columns) == ['row_id', 'responder_6']\n\n    # Ensure the number of rows in the prediction matches the test data\n    assert len(predictions) == len(test)\n    \n    return predictions\n```",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3056980": "Could please help me identify what's wrong with my code? I've tried over 30 times but continue to encounter a \"Submission Format Error.\"\n\n```python\nmodel = create_model()\n\nmodel_paths = [\n    \"/kaggle/input/jane-street/other/default/5/model_0.weights.h5\",\n    \"/kaggle/input/jane-street/other/default/5/model_1.weights.h5\",\n    \"/kaggle/input/jane-street/other/default/5/model_2.weights.h5\",\n    \"/kaggle/input/jane-street/other/default/5/model_3.weights.h5\",\n    \"/kaggle/input/jane-street/other/default/5/model_4.weights.h5\"\n]\n```\n```python\nfrom sklearn.preprocessing import StandardScaler\n\nlags_ : pl.DataFrame | None = None\n\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction.\"\"\"\n    # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n    # Use them as extra features, if you like.\n    global lags_\n    if lags is not None:\n        lags_ = lags\n    \n    test = test.fill_null(0)\n    test = test.join(lags_, on=[\"date_id\", \"time_id\", \"symbol_id\"], how=\"inner\")\n\n    predictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6')\n    )\n\n    feature_input = test.select(\n    ['symbol_id', 'weight', 'feature_00', 'feature_01', 'feature_02', 'feature_03', 'feature_04', 'feature_05', 'feature_06', 'feature_07', 'feature_08', 'feature_09', 'feature_10', 'feature_11', 'feature_12', 'feature_13', 'feature_14', 'feature_15', 'feature_16', 'feature_17', 'feature_18', 'feature_19', 'feature_20', 'feature_21', 'feature_22', 'feature_23', 'feature_24', 'feature_25', 'feature_26', 'feature_27', 'feature_28', 'feature_29', 'feature_30', 'feature_31', 'feature_32', 'feature_33', 'feature_34', 'feature_35', 'feature_36', 'feature_37', 'feature_38', 'feature_39', 'feature_40', 'feature_41', 'feature_42', 'feature_43', 'feature_44', 'feature_45', 'feature_46', 'feature_47', 'feature_48', 'feature_49', 'feature_50', 'feature_51', 'feature_52', 'feature_53', 'feature_54', 'feature_55', 'feature_56', 'feature_57', 'feature_58', 'feature_59', 'feature_60', 'feature_61', 'feature_62', 'feature_63', 'feature_64', 'feature_65', 'feature_66', 'feature_67', 'feature_68', 'feature_69', 'feature_70', 'feature_71', 'feature_72', 'feature_73', 'feature_74', 'feature_75', 'feature_76', 'feature_77', 'feature_78', 'responder_0_lag_1', 'responder_1_lag_1', 'responder_2_lag_1', 'responder_3_lag_1', 'responder_4_lag_1', 'responder_5_lag_1']\n    )\n\n    preds = []\n    for model_path in model_paths:\n        model.load_weights(model_path)  \n        prediction = model.predict(feature_input)  \n        preds.append(prediction)\n    \n    final_pred = sum(preds) / len(preds)\n    \n\n    final_pred_df = pd.DataFrame(final_pred, columns=[\"predictions\"])\n    test_df = pd.DataFrame(test[\"row_id\"], columns=[\"row_id\"])\n    \n    predictions = pd.DataFrame({\n    \"row_id\": test_df[\"row_id\"].values,  \n    \"responder_6\": final_pred_df[\"predictions\"].values  \n    })\n    \n  \n    # Ensure the prediction function returns a DataFrame\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n\n    # Ensure the returned DataFrame has columns 'row_id' and 'responder_6'\n    assert list(predictions.columns) == ['row_id', 'responder_6']\n\n    # Ensure the number of rows in the prediction matches the test data\n    assert len(predictions) == len(test)\n    \n    return predictions\n```"
  }
}