{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Model submission template","metadata":{}},{"cell_type":"markdown","source":">Please <b><span style=\"color:red\">up-vote</span></b> if you find it useful.<br>\n<span style=\"color:red\">Thank you</span>","metadata":{}},{"cell_type":"markdown","source":"^^Feel free to for it and try in your account**","metadata":{}},{"cell_type":"markdown","source":"## Introduction\n\nThis notebook serves as a template for creating a model submission for the [Jane Street Real-Time Market Data Forecasting](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting) competition on Kaggle. It is designed to run both on Kaggle's platform and on a local Linux-based machine. The notebook includes instructions and code to set up the environment, handle data, and make predictions.\n\nIt is based on:\nhttps://www.kaggle.com/code/ryanholbrook/jane-street-rmf-demo-submission\n\n**Note**: The current implementation of the `predict` function generates random predictions for demonstration purposes. You are encouraged to replace it with your actual model.\n\n## Setup and Requirements\n\nTo run this notebook locally on your Linux-based machine, follow the steps below.\n\n### 1. Install Dependencies\n\nEnsure you have the necessary Python packages installed:\n\n```bash\npip install numpy pandas polars grpcio\n```\n\n### 2. Get the Data\n\nDownload the competition data using the Kaggle API:\n\n```bash\nkaggle competitions download -c jane-street-real-time-market-data-forecasting\nunzip jane-street-real-time-market-data-forecasting.zip -d data/\n```\n\n### 3. Create Symlinks\n\nCreate symlinks to mimic the Kaggle environment's directory structure. This allows the code to run both on Kaggle and locally.\n\n```bash\nmkdir -p /kaggle/input/jane-street-real-time-market-data-forecasting\nln -s $(pwd)/data/* /kaggle/input/jane-street-real-time-market-data-forecasting/\n```\n\n### 4. Using Jupyter Notebook\n\nInstall Jupyter Notebook if you haven't already:\n\n```bash\npip install notebook ipykernel jupyter\n```\n\nThen, start the notebook server:\n\n```bash\njupyter notebook\n```","metadata":{}},{"cell_type":"markdown","source":"## Submission requirements","metadata":{}},{"cell_type":"markdown","source":"- Your submission must include a `predict` function that takes in a test DataFrame and returns predictions in the required format.\n- The predictions should be in a DataFrame with columns `['row_id', 'responder_6']`.","metadata":{}},{"cell_type":"code","source":"import os\nimport sys\nimport traceback\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\n\ntry:\n    import kaggle_evaluation.jane_street_inference_server\nexcept ImportError:\n    # Attempt to add the symlinked /kaggle/input/... directory to sys.path\n    input_dir = '/kaggle/input/jane-street-real-time-market-data-forecasting'\n    if os.path.exists(input_dir):\n        sys.path.append(input_dir)\n        print(os.listdir(input_dir))\n        print(os.listdir(os.path.join(input_dir, \"kaggle_evaluation\")))\n        try:\n            import kaggle_evaluation.jane_street_inference_server\n            print(\"Imported local kaggle_evaluation\")\n        except Exception as e:\n            print(\"An error occurred during import:\")\n            traceback.print_exc()\n            print(\"Please analyse the traceback as you may not have installed all dependencies\")\n    else:\n        print(\"Input directory not found. Check environment.\")","metadata":{"execution":{"iopub.status.busy":"2024-10-30T10:20:40.600163Z","iopub.execute_input":"2024-10-30T10:20:40.600566Z","iopub.status.idle":"2024-10-30T10:20:42.330528Z","shell.execute_reply.started":"2024-10-30T10:20:40.600527Z","shell.execute_reply":"2024-10-30T10:20:42.329253Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nfrom pathlib import Path\n\nimport pandas as pd\nimport numpy as np\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import StandardScaler\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.model_selection import train_test_split, TimeSeriesSplit\nfrom sklearn.preprocessing import StandardScaler\nfrom xgboost import XGBRegressor\nfrom sklearn.metrics import r2_score, mean_absolute_error, mean_squared_error\n\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2024-10-30T10:20:42.332546Z","iopub.execute_input":"2024-10-30T10:20:42.333159Z","iopub.status.idle":"2024-10-30T10:20:44.869054Z","shell.execute_reply.started":"2024-10-30T10:20:42.333116Z","shell.execute_reply":"2024-10-30T10:20:44.867956Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Paths","metadata":{}},{"cell_type":"code","source":"INPUT_DIR = Path('/kaggle/input/jane-street-real-time-market-data-forecasting')","metadata":{"execution":{"iopub.status.busy":"2024-10-30T10:20:44.870733Z","iopub.execute_input":"2024-10-30T10:20:44.871283Z","iopub.status.idle":"2024-10-30T10:20:44.877459Z","shell.execute_reply.started":"2024-10-30T10:20:44.871246Z","shell.execute_reply":"2024-10-30T10:20:44.876071Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data reading","metadata":{}},{"cell_type":"markdown","source":"For demonstration purposes, we'll read a subset of the training data.","metadata":{}},{"cell_type":"code","source":"file_name = 'train.parquet/partition_id=0/part-0.parquet'\n\ndf = pd.read_parquet(os.path.join(INPUT_DIR, file_name))\nprint(len(df))\ndf.head(10)","metadata":{"execution":{"iopub.status.busy":"2024-10-30T10:20:44.879754Z","iopub.execute_input":"2024-10-30T10:20:44.880224Z","iopub.status.idle":"2024-10-30T10:20:48.487932Z","shell.execute_reply.started":"2024-10-30T10:20:44.880170Z","shell.execute_reply":"2024-10-30T10:20:48.486897Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n>**Note**: The `predict` function currently generates random predictions for demonstration purposes. You should replace this function with your actual model's prediction logic.","metadata":{}},{"cell_type":"code","source":"lags_: pl.DataFrame | None = None\n\n    \n# THIS IS THE FUNCTION WHICH NEEDS TO BE UPDATED\ndef predict(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n    \"\"\"Make a prediction using a mock model that returns random outputs.\"\"\"\n    global lags_\n    if lags is not None:\n        lags_ = lags\n\n    # Generate random predictions between 0 and 1\n    num_rows = len(test)\n    random_predictions = np.random.rand(num_rows)\n\n    # Create the predictions DataFrame\n    predictions = pl.DataFrame({\n        'row_id': test['row_id'],\n        'responder_6': random_predictions\n    })\n\n    # Ensure the predictions meet the required format\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n    assert predictions.columns == ['row_id', 'responder_6']\n    assert len(predictions) == len(test)\n\n    return predictions\n","metadata":{"execution":{"iopub.status.busy":"2024-10-30T10:20:48.489309Z","iopub.execute_input":"2024-10-30T10:20:48.489764Z","iopub.status.idle":"2024-10-30T10:20:48.497177Z","shell.execute_reply.started":"2024-10-30T10:20:48.489715Z","shell.execute_reply":"2024-10-30T10:20:48.496170Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Test the Predict Function Locally","metadata":{}},{"cell_type":"code","source":"# Simulate a test batch\ntest_batch = pl.DataFrame({\n    'row_id': [1, 2, 3, 4, 5],\n    # Include other necessary feature columns if needed\n})\n# Simulate lags data (optional)\nlags_data = None  # or provide a DataFrame with lag features\n# Call the predict function\npredictions = predict(test_batch, lags_data)\n\nprint(predictions)","metadata":{"execution":{"iopub.status.busy":"2024-10-30T10:20:48.498331Z","iopub.execute_input":"2024-10-30T10:20:48.498718Z","iopub.status.idle":"2024-10-30T10:20:48.573322Z","shell.execute_reply.started":"2024-10-30T10:20:48.498678Z","shell.execute_reply":"2024-10-30T10:20:48.572291Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nprint('KAGGLE_IS_COMPETITION_RERUN', os.getenv('KAGGLE_IS_COMPETITION_RERUN'))\nprint(dir(kaggle_evaluation.jane_street_inference_server))\n\ninference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n        )\n    )\n\n","metadata":{"execution":{"iopub.status.busy":"2024-10-30T10:20:48.574798Z","iopub.execute_input":"2024-10-30T10:20:48.575140Z","iopub.status.idle":"2024-10-30T10:20:48.866338Z","shell.execute_reply.started":"2024-10-30T10:20:48.575106Z","shell.execute_reply":"2024-10-30T10:20:48.865370Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n---\n\n## How to Use This Notebook\n\nTo use this notebook on your local machine:\n\n1. **Set Up the Environment**: Follow the instructions in the \"Setup and Requirements\" section to install dependencies and set up the data.\n\n2. **Understand the Code Structure**:\n   - The notebook imports necessary libraries and handles the environment setup to ensure compatibility between Kaggle and local environments.\n   - The `predict` function is where you should implement your model's prediction logic.\n   - The inference server is set up to run the model either in the Kaggle competition environment or locally.\n\n3. **Replace the Mock Model**:\n   - Currently, the `predict` function generates random predictions. Replace the random predictions with your model's outputs.\n   - Ensure that your predictions meet the required format: a DataFrame with columns `['row_id', 'responder_6']`.\n\n4. **Run the Notebook**:\n   - Execute all cells sequentially after making necessary changes.\n   - The notebook will simulate predictions and, if not in Kaggle's competition rerun environment, will run the inference server locally.\n\n---\n\n## Conclusion\n\nThis notebook provides a template for submitting your model to the Jane Street competition. By following the instructions and replacing the mock `predict` function with your actual model, you can create a valid submission that runs both locally and on Kaggle.\n\n---","metadata":{}}]}