{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":84493,"databundleVersionId":9871156,"sourceType":"competition"},{"sourceId":9625192,"sourceType":"datasetVersion","datasetId":5875295},{"sourceId":9631435,"sourceType":"datasetVersion","datasetId":5879957},{"sourceId":9640394,"sourceType":"datasetVersion","datasetId":5882430},{"sourceId":201255000,"sourceType":"kernelVersion"},{"sourceId":201377683,"sourceType":"kernelVersion"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"One of the outstanding members of the Kaggle community, grandmaster, gives a link to his [***work***](https://www.kaggle.com/code/cdeotte/top-solutions-ensemble-0-947), where he shows the refinement of the prediction using an experimental example. In order to try to refine our predictions at the [Jane Street Real-Time Market Data Forecasting](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/code?competitionId=84493&sortBy=scoreDescending&excludeNonAccessedDatasources=true) competition.\n\nThis approach has proven itself in the following competitions: 1. [ISIC 2024 - Skin Cancer Detection with 3D-TBP](https://www.kaggle.com/competitions/isic-2024-challenge/code?competitionId=63056&sortBy=scoreDescending&excludeNonAccessedDatasources=true), 2. [RSNA 2024 Lumbar Spine Degenerative Classification](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/code?competitionId=71549&sortBy=scoreAscending&excludeNonAccessedDatasources=true), 3. [NeurIPS - Ariel Data Challenge 2024](https://www.kaggle.com/competitions/ariel-data-challenge-2024/code?competitionId=70367&sortBy=scoreDescending&excludeNonAccessedDatasources=true), 4. [Child Mind Institute — Problematic Internet Use](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/code?competitionId=81933&sortBy=scoreDescending&excludeNonAccessedDatasources=true), 5. [BrisT1D Blood Glucose Prediction Competition](https://www.kaggle.com/competitions/brist1d/code?competitionId=82611&sortBy=scoreAscending&excludeNonAccessedDatasources=true), 6. [Eedi - Mining Misconceptions in Mathematics](https://www.kaggle.com/competitions/eedi-mining-misconceptions-in-mathematics/code?competitionId=82695&sortBy=scoreDescending&excludeNonAccessedDatasources=true), 7. [Connect X](https://www.kaggle.com/competitions/connectx/leaderboard?), 8. [Loan Approval Prediction {PS-S4.E10}](https://www.kaggle.com/competitions/brist1d/code)\n\nAnd accordingly in the following notebooks: 1. [ISIC | Ensemble of solutions](https://www.kaggle.com/code/vyacheslavbolotin/isic-2024-ensemble-of-solutions), 2. [RSNA | Ensemble of solutions](https://www.kaggle.com/code/vyacheslavbolotin/rsna-ensemble-of-solutions), 3. [Ariel | Ensemble of solutions](https://www.kaggle.com/code/vyacheslavbolotin/ariel-ensemble-of-solutions), 4. [CMI | Ensemble of solutions](https://www.kaggle.com/code/vyacheslavbolotin/cmi-ensemble-of-solutions), 5. [BrisT1D | Ensemble of solutions](https://www.kaggle.com/code/vyacheslavbolotin/brist1d-ensemble-of-solutions), 6. [Eedi | Ensemble of solutions](https://www.kaggle.com/code/vyacheslavbolotin/eedi-ensemble-of-solutions), 7. [agents Connect X](https://www.kaggle.com/code/vyacheslavbolotin/agents-connect-x), 8. [PS-S4.E10 | Ensemble of solutions](https://www.kaggle.com/code/vyacheslavbolotin/pss4e10-ensemble-of-solutions) \n\n\n#### Jane Street | Ensemble  of solutions:\n\n1. [0.0044](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/code?competitionId=84493&sortBy=scoreDescending&excludeNonAccessedDatasources=true) UAE [JaneStreet2024|Baseline|Submission|V1](https://www.kaggle.com/code/ravi20076/janestreet2024-baseline-submission-v1) by grandmaster [Ravi Ramakrishnan](https://www.kaggle.com/ravi20076)\n2. [0.0040](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/code?competitionId=84493&sortBy=scoreDescending&excludeNonAccessedDatasources=true) China [🥇🥇Jane Street Baseline lgb, xgb and catboost🥇🥇](https://www.kaggle.com/code/yuanzhezhou/jane-street-baseline-lgb-xgb-and-catboost) by grandmaster [yuanzhe zhou](https://www.kaggle.com/yuanzhezhou)\n3. [0.0034](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/code?competitionId=84493&sortBy=scoreDescending&excludeNonAccessedDatasources=true) South Korea [JS - train & infer lgbm & xgb & cat w/ custom eval](https://www.kaggle.com/code/cy4ego/js-train-infer-lgbm-xgb-cat-w-custom-eval) by expert [kcy4](https://www.kaggle.com/cy4ego)\n4. [0.0023](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/code?competitionId=84493&sortBy=scoreDescending&excludeNonAccessedDatasources=true) Australia [Jane Street:Cat_base](https://www.kaggle.com/code/backpaker/jane-street-cat-base) by expert [Backpacker](https://www.kaggle.com/backpaker)\n5. [0.0026](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/code?competitionId=84493&sortBy=scoreDescending&excludeNonAccessedDatasources=true) China [JS Ridge baseline](https://www.kaggle.com/code/yunsuxiaozi/js-ridge-baseline) by master [yunsuxiaozi](https://www.kaggle.com/yunsuxiaozi)\n\n\noptions\n- option 1 -> V08 solutions.(4,4), random.choise[solutions], Lb=0.0023\n- option 2 -> V11 solutions.(1,1), random.choise[solutions], Lb=0.0044\n- option 3 -> V29 solutions.(3,3), random.choise[solutions], Lb=0.0034\n- option 4 -> V33 solutions.(2,2), random.choise[solutions], Lb=0.0040\n\ncurrent options\n- option 5 -> V40 solutions.(5,5), random.choise[solutions], Lb=0.0026\n\nnext options\n- option 7 -> solutions.(3,5), random.choise[solutions], Lb=**?**\n- option 6 -> solutions.(2,5), random.choise[solutions], Lb=**?**\n- option 8 -> solutions.(4,5), random.choise[solutions], Lb=**?**","metadata":{}},{"cell_type":"markdown","source":"At the end of the competition, all data will be grouped and presented in the usual tables. [Stat | Ensemble of solutions](https://www.kaggle.com/code/vyacheslavbolotin/stat-ensemble-of-solutions)","metadata":{}},{"cell_type":"code","source":"OPTION,ENSEMBLE_SOLUTIONS = 'option 5',['SOLUTION_1','SOLUTION_2','SOLUTION_3','SOLUTION_4','SOLUTION_5']\n#OPTION,ENSEMBLE_SOLUTIONS = 'option 7',['SOLUTION_3','SOLUTION_5']","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2024-10-22T11:13:15.633137Z","iopub.execute_input":"2024-10-22T11:13:15.633555Z","iopub.status.idle":"2024-10-22T11:13:15.639198Z","shell.execute_reply.started":"2024-10-22T11:13:15.633514Z","shell.execute_reply":"2024-10-22T11:13:15.637514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1. [JaneStreet2024|Baseline|Submission|V1](https://www.kaggle.com/code/ravi20076/janestreet2024-baseline-submission-v1) Lb=0.0044\n### [Ravi Ramakrishnan](https://www.kaggle.com/ravi20076)","metadata":{}},{"cell_type":"code","source":"if 'SOLUTION_1' in ENSEMBLE_SOLUTIONS:\n    \n    !pip install polars[gpu]==1.9.0 -q --no-index --find-links=/kaggle/input/janestreet2024-imports-v1/polars\n    !pip install lightgbm==4.5.0 -q --no-index --find-links=/kaggle/input/janestreet2024-imports-v1/packages\n    !pip install scikit-learn==1.5.2 -q --no-index --find-links=/kaggle/input/janestreet2024-imports-v1/packages\n\n    exec(\n        open(\"/kaggle/input/janestreet2024-imports-v1/myimports.py\", \"r\"\n            ).read()\n    )\n\n    print()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:13:15.640765Z","iopub.execute_input":"2024-10-22T11:13:15.641074Z","iopub.status.idle":"2024-10-22T11:13:50.852213Z","shell.execute_reply.started":"2024-10-22T11:13:15.641042Z","shell.execute_reply":"2024-10-22T11:13:50.851162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_1' in ENSEMBLE_SOLUTIONS:\n\n    target     = \"responder_6\"\n    op_path    = f\"/kaggle/working\"\n    ip_path    = f\"/kaggle/input/janestreet2024-dataload-v1\"\n    state      = 42\n    method     = \"CB1R\"","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:13:50.853708Z","iopub.execute_input":"2024-10-22T11:13:50.854077Z","iopub.status.idle":"2024-10-22T11:13:50.859872Z","shell.execute_reply.started":"2024-10-22T11:13:50.854041Z","shell.execute_reply":"2024-10-22T11:13:50.858964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_1' in ENSEMBLE_SOLUTIONS:\n\n    sel_cols = \\\n    [\n    'symbol_id', \n    'feature_00', 'feature_01', 'feature_02', 'feature_03', 'feature_04',\n    'feature_05', 'feature_06', 'feature_07', 'feature_08', 'feature_09', 'feature_10',\n    'feature_11', 'feature_12', 'feature_13', 'feature_14', 'feature_15', 'feature_16',\n    'feature_17', 'feature_18', 'feature_19', 'feature_20', 'feature_21', 'feature_22',\n    'feature_23', 'feature_24', 'feature_25', 'feature_26', 'feature_27', 'feature_28',\n    'feature_29', 'feature_30', 'feature_31', 'feature_32', 'feature_33', 'feature_34',\n    'feature_35', 'feature_36', 'feature_37', 'feature_38', 'feature_39', 'feature_40',\n    'feature_41', 'feature_42', 'feature_43', 'feature_44', 'feature_45', 'feature_46',\n    'feature_47', 'feature_48', 'feature_49', 'feature_50', 'feature_51', 'feature_52',\n    'feature_53', 'feature_54', 'feature_55', 'feature_56', 'feature_57', 'feature_58',\n    'feature_59', 'feature_60', 'feature_61', 'feature_62', 'feature_63', 'feature_64',\n    'feature_65', 'feature_66', 'feature_67', 'feature_68', 'feature_69', 'feature_70',\n    'feature_71', 'feature_72', 'feature_73', 'feature_74', 'feature_75', 'feature_76',\n    'feature_77', 'feature_78'\n    ]\n\n    all_files  = sorted(os.listdir(f\"/kaggle/input/janestreetpublicv1\"))\n    sel_models = [\"CBV1_2.joblib\", \"CBV1_3.joblib\", \"CBV1_5.joblib\", \n                  \"LGBMV1_1.joblib\", \"LGBMV1_2.joblib\", \"LGBMV1_4.joblib\", \"LGBMV1_5.joblib\",\n                 ]\n    all_files  = list(set(all_files).intersection(set(sel_models)))\n\n    models = []\n    for file in sorted(all_files):\n        PrintColor(f\"---> Current model file - {file}\", color = Fore.CYAN)\n        fitted_model = \\\n        joblib.load(\n            os.path.join(f\"/kaggle/input/janestreetpublicv1\", file)\n        )[\"Online\"]\n\n        models.append(fitted_model)\n        del fitted_model\n\n    PrintColor(f\"\\n---> Models for inference\\n\")\n    pprint(models)\n\n    print()\n    collect();","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:13:50.861956Z","iopub.execute_input":"2024-10-22T11:13:50.862290Z","iopub.status.idle":"2024-10-22T11:14:16.479132Z","shell.execute_reply.started":"2024-10-22T11:13:50.862257Z","shell.execute_reply":"2024-10-22T11:14:16.478330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_1' in ENSEMBLE_SOLUTIONS:\n\n    lags_ : pl.DataFrame | None = None\n\n    def predict_1(\n        test: pl.DataFrame, \n        lags: pl.DataFrame | None\n    ) -> pl.DataFrame | pd.DataFrame:\n        \"This is the inference and submission function used to predict the test set for the competition\"\n\n#         lags_, models, sel_cols, target\n\n        if lags is not None:\n            lags_ = lags\n\n        test_preds = []\n        for model in tqdm(models):\n            test_preds.append(\n                model.predict(\n                    test.select(pl.col(sel_cols)).to_pandas()\n                )\n            )\n\n        test_preds = \\\n        np.average(\n            np.stack(test_preds, axis=1), \n            axis    = 1,\n            weights = [0.10, 0.10, 0.10, 0.10, 0.25, 0.25, 0.10]\n        )\n\n        predictions = \\\n        test.select('row_id').\\\n        with_columns(\n            pl.Series(\n                name   = 'responder_6', \n                values = np.clip(test_preds, a_min = -5, a_max = 5),\n                dtype  = pl.Float64,\n            )\n        )  \n\n        assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n        assert predictions.columns == ['row_id', 'responder_6']\n        assert len(predictions) == len(test)\n        return predictions","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:16.480466Z","iopub.execute_input":"2024-10-22T11:14:16.481030Z","iopub.status.idle":"2024-10-22T11:14:16.488949Z","shell.execute_reply.started":"2024-10-22T11:14:16.480988Z","shell.execute_reply":"2024-10-22T11:14:16.488305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. [🥇🥇Jane Street Baseline lgb, xgb and catboost🥇🥇](https://www.kaggle.com/code/yuanzhezhou/jane-street-baseline-lgb-xgb-and-catboost) Lb=0.0040\n### [yuanzhe zhou](https://www.kaggle.com/yuanzhezhou)","metadata":{}},{"cell_type":"code","source":"if 'SOLUTION_2' in ENSEMBLE_SOLUTIONS:\n    \n    import os\n    import joblib \n    import random\n    import pandas as pd\n    import polars as pl\n    import lightgbm as lgb\n    import xgboost as xgb\n    import catboost as cbt\n    import numpy as np \n\n    from joblib import Parallel, delayed\n    \n    import kaggle_evaluation.jane_street_inference_server\n\n    # !pip install lightgbm==4.2.0 -i https://mirrors.aliyun.com/pypi/simple/\n    # !pip install catboost==1.2.7 -i https://mirrors.aliyun.com/pypi/simple/\n    # !pip install xgboost==2.0.3 -i https://mirrors.aliyun.com/pypi/simple/\n    # !pip install joblib==1.4.2 -i https://mirrors.aliyun.com/pypi/simple/\n\n\n    def reduce_mem_usage(self, float16_as32=True):\n        #memory_usage()是df每列的内存使用量,sum是对它们求和, B->KB->MB\n        start_mem = df.memory_usage().sum() / 1024**2\n        print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n\n        for col in df.columns:#遍历每列的列名\n            col_type = df[col].dtype#列名的type\n            if col_type != object and str(col_type)!='category':#不是object也就是说这里处理的是数值类型的变量\n                c_min,c_max = df[col].min(),df[col].max() #求出这列的最大值和最小值\n                if str(col_type)[:3] == 'int':#如果是int类型的变量,不管是int8,int16,int32还是int64\n                    #如果这列的取值范围是在int8的取值范围内,那就对类型进行转换 (-128 到 127)\n                    if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                        df[col] = df[col].astype(np.int8)\n                    #如果这列的取值范围是在int16的取值范围内,那就对类型进行转换(-32,768 到 32,767)\n                    elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                        df[col] = df[col].astype(np.int16)\n                    #如果这列的取值范围是在int32的取值范围内,那就对类型进行转换(-2,147,483,648到2,147,483,647)\n                    elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                        df[col] = df[col].astype(np.int32)\n                    #如果这列的取值范围是在int64的取值范围内,那就对类型进行转换(-9,223,372,036,854,775,808到9,223,372,036,854,775,807)\n                    elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                        df[col] = df[col].astype(np.int64)  \n                else:#如果是浮点数类型.\n                    #如果数值在float16的取值范围内,如果觉得需要更高精度可以考虑float32\n                    if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                        if float16_as32:#如果数据需要更高的精度可以选择float32\n                            df[col] = df[col].astype(np.float32)\n                        else:\n                            df[col] = df[col].astype(np.float16)  \n                    #如果数值在float32的取值范围内，对它进行类型转换\n                    elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                        df[col] = df[col].astype(np.float32)\n                    #如果数值在float64的取值范围内，对它进行类型转换\n                    else:\n                        df[col] = df[col].astype(np.float64)\n        #计算一下结束后的内存\n        end_mem = df.memory_usage().sum() / 1024**2\n        print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n        #相比一开始的内存减少了百分之多少\n        print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n\n        return df","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:16.490307Z","iopub.execute_input":"2024-10-22T11:14:16.490630Z","iopub.status.idle":"2024-10-22T11:14:16.510871Z","shell.execute_reply.started":"2024-10-22T11:14:16.490597Z","shell.execute_reply":"2024-10-22T11:14:16.510034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_2' in ENSEMBLE_SOLUTIONS:\n    \n    # Define the path to the input data directory\n    # If the local directory exists, use it; otherwise, use the Kaggle input directory\n    input_path = './jane-street-real-time-market-data-forecasting/' if os.path.exists('./jane-street-real-time-market-data-forecasting') else '/kaggle/input/jane-street-real-time-market-data-forecasting/'\n\n    # Flag to determine if the script is in training mode or not\n    TRAINING = False\n\n    # Define the feature names based on the number of features (79 in this case)\n    feature_names = [f\"feature_{i:02d}\" for i in range(79)]\n\n    # Number of validation dates to use\n    num_valid_dates = 100\n\n    # Number of dates to skip from the beginning of the dataset\n    skip_dates = 500\n\n    # Number of folds for cross-validation\n    N_fold = 5\n\n    # If in training mode, load the training data\n    if TRAINING:\n        # Load the training data from a Parquet file\n        df = pd.read_parquet(f'{input_path}/train.parquet')\n\n        # Reduce memory usage of the DataFrame (function not provided here)\n        df = reduce_mem_usage(df, False)\n\n        # Filter the DataFrame to include only dates greater than or equal to skip_dates\n        df = df[df['date_id'] >= skip_dates].reset_index(drop=True)\n\n        # Get unique dates from the DataFrame\n        dates = df['date_id'].unique()\n\n        # Define validation dates as the last `num_valid_dates` dates\n        valid_dates = dates[-num_valid_dates:]\n\n        # Define training dates as all dates except the last `num_valid_dates` dates\n        train_dates = dates[:-num_valid_dates]\n\n        # Display the last few rows of the DataFrame (for debugging purposes)\n        print(df.tail())","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:16.512161Z","iopub.execute_input":"2024-10-22T11:14:16.512447Z","iopub.status.idle":"2024-10-22T11:14:16.526908Z","shell.execute_reply.started":"2024-10-22T11:14:16.512417Z","shell.execute_reply":"2024-10-22T11:14:16.526045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_2' in ENSEMBLE_SOLUTIONS:\n    \n    # Create a directory to store the trained models\n    os.system('mkdir models')\n\n    # Define the path to load pre-trained models (if not in training mode)\n    model_path = '/kaggle/input/jsbaselinezyz'\n\n    # If in training mode, prepare validation data\n    if TRAINING:\n        # Extract features, target, and weights for validation dates\n        X_valid = df[feature_names].loc[df['date_id'].isin(valid_dates)]\n        y_valid = df['responder_6'].loc[df['date_id'].isin(valid_dates)]\n        w_valid = df['weight'].loc[df['date_id'].isin(valid_dates)]\n\n    # Initialize a list to store trained models\n    models = []\n\n    # Function to train a model or load a pre-trained model\n    def train(model_dict, model_name='lgb'):\n        if TRAINING:\n            # Select dates for training based on the fold number\n            selected_dates = [date for ii, date in enumerate(train_dates) if ii % N_fold != i]\n\n            # Get the model from the dictionary\n            model = model_dict[model_name]\n\n            # Extract features, target, and weights for the selected training dates\n            X_train = df[feature_names].loc[df['date_id'].isin(selected_dates)]\n            y_train = df['responder_6'].loc[df['date_id'].isin(selected_dates)]\n            w_train = df['weight'].loc[df['date_id'].isin(selected_dates)]\n\n            # Train the model based on the type (LightGBM, XGBoost, or CatBoost)\n            if model_name == 'lgb':\n                # Train LightGBM model with early stopping and evaluation logging\n                model.fit(X_train, y_train, w_train,  \n                          eval_metric=[r2_lgb],\n                          eval_set=[(X_valid, y_valid, w_valid)], \n                          callbacks=[\n                              lgb.early_stopping(100), \n                              lgb.log_evaluation(10)\n                          ])\n\n            elif model_name == 'cbt':\n                # Prepare evaluation set for CatBoost\n                evalset = cbt.Pool(X_valid, y_valid, weight=w_valid)\n\n                # Train CatBoost model with early stopping and verbose logging\n                model.fit(X_train, y_train, sample_weight=w_train, \n                          eval_set=[evalset], \n                          verbose=10, \n                          early_stopping_rounds=100)\n\n            else:\n                # Train XGBoost model with early stopping and verbose logging\n                model.fit(X_train, y_train, sample_weight=w_train, \n                          eval_set=[(X_valid, y_valid)], \n                          sample_weight_eval_set=[w_valid], \n                          verbose=10, \n                          early_stopping_rounds=100)\n\n            # Append the trained model to the list\n            models.append(model)\n\n            # Save the trained model to a file\n            joblib.dump(model, f'./models/{model_name}_{i}.model')\n\n            # Delete training data to free up memory\n            del X_train\n            del y_train\n            del w_train\n\n            # Collect garbage to free up memory\n            import gc\n            gc.collect()\n\n        else:\n            # If not in training mode, load the pre-trained model from the specified path\n            models.append(joblib.load(f'{model_path}/{model_name}_{i}.model'))\n\n        return \n\n    # Custom R2 metric for XGBoost\n    def r2_xgb(y_true, y_pred, sample_weight):\n        r2 = 1 - np.average((y_pred - y_true) ** 2, weights=sample_weight) / (np.average((y_true) ** 2, weights=sample_weight) + 1e-38)\n        return -r2\n\n    # Custom R2 metric for LightGBM\n    def r2_lgb(y_true, y_pred, sample_weight):\n        r2 = 1 - np.average((y_pred - y_true) ** 2, weights=sample_weight) / (np.average((y_true) ** 2, weights=sample_weight) + 1e-38)\n        return 'r2', r2, True\n\n    # Custom R2 metric for CatBoost\n    class r2_cbt(object):\n        def get_final_error(self, error, weight):\n            return 1 - error / (weight + 1e-38)\n\n        def is_max_optimal(self):\n            return True\n\n        def evaluate(self, approxes, target, weight):\n            assert len(approxes) == 1\n            assert len(target) == len(approxes[0])\n\n            approx = approxes[0]\n\n            error_sum = 0.0\n            weight_sum = 0.0\n\n            for i in range(len(approx)):\n                w = 1.0 if weight is None else weight[i]\n                weight_sum += w * (target[i] ** 2)\n                error_sum += w * ((approx[i] - target[i]) ** 2)\n\n            return error_sum, weight_sum\n\n    # Dictionary to store different models with their configurations\n    model_dict = {\n        'lgb': lgb.LGBMRegressor(n_estimators=500, device='gpu', gpu_use_dp=True, objective='l2'),\n        'xgb': xgb.XGBRegressor(n_estimators=2000, learning_rate=0.1, max_depth=6, tree_method='hist', device=\"cuda\", objective='reg:squarederror', eval_metric=r2_xgb, disable_default_eval_metric=True),\n        'cbt': cbt.CatBoostRegressor(iterations=1000, learning_rate=0.05, task_type='GPU', loss_function='RMSE', eval_metric=r2_cbt()),\n    }\n\n    # Train models for each fold\n    for i in range(N_fold):\n        train(model_dict, 'lgb')\n        train(model_dict, 'xgb')\n        train(model_dict, 'cbt')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"scrolled":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:16.528187Z","iopub.execute_input":"2024-10-22T11:14:16.528991Z","iopub.status.idle":"2024-10-22T11:14:51.698393Z","shell.execute_reply.started":"2024-10-22T11:14:16.528945Z","shell.execute_reply":"2024-10-22T11:14:51.697350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_2' in ENSEMBLE_SOLUTIONS:    \n    \n    lags_ : pl.DataFrame | None = None\n\n    # Replace this function with your inference code.\n    # You can return either a Pandas or Polars dataframe, though Polars is recommended.\n    # Each batch of predictions (except the very first) must be returned within 10 minutes of the batch features being provided.\n    def predict_2(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n        \"\"\"Make a prediction.\"\"\"\n        # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n        # Use them as extra features, if you like.\n        global lags_\n        if lags is not None:\n            lags_ = lags\n\n        predictions = test.select(\n            'row_id',\n            pl.lit(0.0).alias('responder_6'),\n        )\n\n        feat = test[feature_names].to_numpy()\n\n        pred = [model.predict(feat) for model in models]\n        pred = np.mean(pred, axis=0)\n\n        predictions = predictions.with_columns(pl.Series('responder_6', pred.ravel()))\n\n        # The predict function must return a DataFrame\n        assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n        # with columns 'row_id', 'responer_6'\n        assert list(predictions.columns) == ['row_id', 'responder_6']\n        # and as many rows as the test data.\n        assert len(predictions) == len(test)\n\n        return predictions","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.701669Z","iopub.execute_input":"2024-10-22T11:14:51.702028Z","iopub.status.idle":"2024-10-22T11:14:51.710615Z","shell.execute_reply.started":"2024-10-22T11:14:51.701992Z","shell.execute_reply":"2024-10-22T11:14:51.709707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When your notebook is run on the hidden test set, inference_server.serve must be called within 15 minutes of the notebook starting or the gateway will throw an error. If you need more than 15 minutes to load your model you can do so during the very first `predict` call, which does not have the usual 10 minute response deadline.","metadata":{"_kg_hide-input":true,"_kg_hide-output":true}},{"cell_type":"markdown","source":"## 3. [JS - train & infer lgbm & xgb & cat w/ custom eval](https://www.kaggle.com/code/cy4ego/js-train-infer-lgbm-xgb-cat-w-custom-eval) Lb=0.0034\n### [kcy4](https://www.kaggle.com/cy4ego)","metadata":{}},{"cell_type":"code","source":"if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:\n    \n    import glob \n    import gc \n    import re\n    import os\n    import sys\n    import random\n    import numpy as np\n    import polars as pl\n    import pandas as pd\n    import lightgbm as lgb\n    import xgboost as xgb\n    import catboost as cat \n    from tqdm import tqdm \n    import seaborn as sns\n    from typing import Tuple\n        # cat_pred = models[2].predict(feat)\n        \n    import kaggle_evaluation.jane_street_inference_server\n\n\n    pd.set_option(\"display.max_rows\", 2000)\n    pd.options.display.float_format = \"{:,.6f}\".format\n\n    SEED = 42\n    def seed_everything(seed=SEED):\n        os.environ[\"PYTHONHASHSEED\"] = str(seed)\n        random.seed(seed)\n        np.random.seed(seed)\n\n    seed_everything()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.711658Z","iopub.execute_input":"2024-10-22T11:14:51.711947Z","iopub.status.idle":"2024-10-22T11:14:51.729963Z","shell.execute_reply.started":"2024-10-22T11:14:51.711902Z","shell.execute_reply":"2024-10-22T11:14:51.729024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:\n    \n    N_ESTIMATORS = 8_000\n    N_SPLITS = 5\n    EARLY_STOP = 100\n    ### DATA ###\n    # Train data is too big to be loaded in the kaggle notebook. Make some adjustments if you want.\n    # date_id > 1100\n    NUM_ROWS_NOT_TO_BE_USED = 25023058 if N_ESTIMATORS > 10 else 39023058\n    NUM_VALID_DATES = 100 if N_ESTIMATORS > 10 else 10\n\n    INPUT_DIR = '/kaggle/input/jane-street-real-time-market-data-forecasting'\n    OUTPUT_DIR = '/kaggle/output'","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.731018Z","iopub.execute_input":"2024-10-22T11:14:51.731281Z","iopub.status.idle":"2024-10-22T11:14:51.745020Z","shell.execute_reply.started":"2024-10-22T11:14:51.731251Z","shell.execute_reply":"2024-10-22T11:14:51.743998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:\n    \n    TARGET = 'responder_6'\n    TIME_COLS = ['date_id', 'time_id']\n    LEAD_COLS = ['symbol_id', 'weight']\n    RESPONDER_COLS = [f\"responder_{i}\" for i in range(9)]\n    FEAT_COLS = [f\"feature_{i:02d}\" for i in range(79)]","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.746024Z","iopub.execute_input":"2024-10-22T11:14:51.746344Z","iopub.status.idle":"2024-10-22T11:14:51.756673Z","shell.execute_reply.started":"2024-10-22T11:14:51.746311Z","shell.execute_reply":"2024-10-22T11:14:51.755813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Custom eval metric functions\n- Made some changes to [yuanzhe zhou](https://www.kaggle.com/code/yuanzhezhou/jane-street-baseline-lgb-xgb-and-catboost)'s eval function","metadata":{"_kg_hide-input":true}},{"cell_type":"code","source":"if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS: \n    \n    # Edited by kcy4\n    # Removed r2_lgb and r2_xgb then merged them into one function.\n    # Changed to use lgb.Dataset or xgb.DMatrix\n    def r2_gbt(y_pred, dtrain: lgb.Dataset|xgb.DMatrix):\n        y_true = dtrain.get_label()\n        weight = dtrain.get_weight()\n        r2 = 1 - np.average((y_pred - y_true) ** 2, weights=weight) / (np.average((y_true) ** 2, weights=weight) + 1e-38)\n        if isinstance(dtrain, lgb.Dataset):\n            return 'r2', r2, True\n        else: # for xgboost\n            return 'r2', -r2\n\n    # No touch\n    # Custom R2 metric for CatBoost\n    class r2_cbt(object):\n        def get_final_error(self, error, weight):\n            return 1 - error / (weight + 1e-38)\n\n        def is_max_optimal(self):\n            return True\n\n        def evaluate(self, approxes, target, weight):\n            assert len(approxes) == 1\n            assert len(target) == len(approxes[0])\n\n            approx = approxes[0]\n\n            error_sum = 0.0\n            weight_sum = 0.0\n\n            for i in range(len(approx)):\n                w = 1.0 if weight is None else weight[i]\n                weight_sum += w * (target[i] ** 2)\n                error_sum += w * ((approx[i] - target[i]) ** 2)\n\n            return error_sum, weight_sum","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.757662Z","iopub.execute_input":"2024-10-22T11:14:51.757945Z","iopub.status.idle":"2024-10-22T11:14:51.768719Z","shell.execute_reply.started":"2024-10-22T11:14:51.757901Z","shell.execute_reply":"2024-10-22T11:14:51.767786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:\n    \n    MODEL_NAMES = ['lgb', 'xgb', 'cat'][:1]\n\n    LGB_PARAMS = {\n        'objective'            : 'l2',\n        'boosting_type'        : 'gbdt',\n        'learning_rate'        : 0.02,\n        'num_leaves'           : 63,\n        'verbose'              : -1,\n        'random_state'         : SEED,\n        # 'device'               : 'gpu',\n        # 'gpu_platform_id'      : 0,\n        # 'gpu_device_id'        : 0,\n    }\n\n    XGB_PARAMS = {\n        'objective'            : 'reg:squarederror',\n        'eval_metric'          : 'rmse',\n        'disable_default_eval_metric': True,\n        # 'device'               : 'cuda:0',\n        'tree_method'          : 'hist',\n        'learning_rate'        : 0.05,\n        # 'eval_metric'          : r2_gbt,\n    }\n\n    CAT_PARAMS={\n        # 'task_type'            : 'GPU',\n        'loss_function'        : 'RMSE',\n        'eval_metric'          : r2_cbt(),\n        'n_estimators'         : N_ESTIMATORS,\n        'learning_rate'        : 0.05,\n        'verbose'              : 0,\n        'random_state'         : SEED,\n        'early_stopping_rounds': 100,\n    }","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.769888Z","iopub.execute_input":"2024-10-22T11:14:51.770297Z","iopub.status.idle":"2024-10-22T11:14:51.783491Z","shell.execute_reply.started":"2024-10-22T11:14:51.770252Z","shell.execute_reply":"2024-10-22T11:14:51.782707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:\n    \n    # time_df = pd.read_parquet(f\"{INPUT_DIR}/train.parquet\", columns=TIME_COLS)[NUM_ROWS_NOT_TO_BE_USED:].astype(np.uint16)\n    # time_df[time_df['date_id'] <= 1100].shape","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.784584Z","iopub.execute_input":"2024-10-22T11:14:51.784839Z","iopub.status.idle":"2024-10-22T11:14:51.796082Z","shell.execute_reply.started":"2024-10-22T11:14:51.784810Z","shell.execute_reply":"2024-10-22T11:14:51.795339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:\n    \n    def get_df():\n        print('Time data loading...')\n        time_df = pd.read_parquet(f\"{INPUT_DIR}/train.parquet\", columns=TIME_COLS)[NUM_ROWS_NOT_TO_BE_USED:].astype(np.uint16)\n        print('Lead data loading...')\n        lead_df = pd.read_parquet(f\"{INPUT_DIR}/train.parquet\", columns=LEAD_COLS)[NUM_ROWS_NOT_TO_BE_USED:]\n        lead_df['symbol_id'] = lead_df['symbol_id'].astype(np.uint32)\n        lead_df['weight'] = lead_df['weight'].astype(np.float32)\n        print('Target data loading...')\n        responder_6_df = pd.read_parquet(f\"{INPUT_DIR}/train.parquet\", columns=[TARGET])[NUM_ROWS_NOT_TO_BE_USED:].astype(np.float32)\n        print('Feature data loading...')\n        feat_dfs = []\n        num_chunk = 10\n        chunk_unit_len = (len(FEAT_COLS) // 10) + 1\n        for i in tqdm(range(10), total=10):\n            read_feat_cols = FEAT_COLS[i*chunk_unit_len:(i+1)*chunk_unit_len]\n            feat_dfs.append(pd.read_parquet(f\"{INPUT_DIR}/train.parquet\", columns=read_feat_cols)[NUM_ROWS_NOT_TO_BE_USED:].astype(np.float32))\n\n        return pd.concat([time_df]+[lead_df]+feat_dfs +[responder_6_df], axis=1)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.797202Z","iopub.execute_input":"2024-10-22T11:14:51.797514Z","iopub.status.idle":"2024-10-22T11:14:51.806896Z","shell.execute_reply.started":"2024-10-22T11:14:51.797465Z","shell.execute_reply":"2024-10-22T11:14:51.806029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:\n    \n    def train_model(df, model_names):\n        models = [None, None, None]\n        dates = df['date_id'].unique()\n        train_dates = dates[:-NUM_VALID_DATES]\n        valid_dates = dates[-NUM_VALID_DATES:]\n\n        tr_weight = df['weight'].loc[df['date_id'].isin(train_dates)]\n        val_weight = df['weight'].loc[df['date_id'].isin(valid_dates)]\n\n        tr_y = df.loc[df['date_id'].isin(train_dates), TARGET]\n        val_y = df.loc[df['date_id'].isin(valid_dates), TARGET]\n\n        tr_X = df.loc[df['date_id'].isin(train_dates)][FEAT_COLS]\n        val_X = df.loc[df['date_id'].isin(valid_dates)][FEAT_COLS]\n\n        if 'lgb' in model_names:\n            print('lgb model training started. It takes some minutes. Have a coffee!')\n            tr_ds = lgb.Dataset(tr_X, label=tr_y, weight=tr_weight)\n            te_ds = lgb.Dataset(val_X, label=val_y, weight=val_weight, reference=tr_ds)\n            model = lgb.train(\n                LGB_PARAMS, tr_ds, N_ESTIMATORS, valid_sets=[te_ds], feval=r2_gbt,\n                callbacks=[lgb.early_stopping(EARLY_STOP)]\n            )\n            print(f\"lgb: {model.best_score['valid_0']['r2']:.06f}\")\n            del tr_ds, te_ds\n            gc.collect()\n            models[0] = model\n        if 'xgb' in model_names:\n            print('xgb model training started. It takes some minutes. Have a coffee!')\n            tr_ds = xgb.DMatrix(tr_X, label=tr_y, weight=tr_weight)\n            te_ds = xgb.DMatrix(val_X, label=val_y, weight=val_weight)\n            model = xgb.train(XGB_PARAMS, tr_ds, N_ESTIMATORS, evals=[(te_ds, 'eval')], early_stopping_rounds=EARLY_STOP, verbose_eval=False, custom_metric=r2_gbt)\n            print(f\"xgb: {-model.best_score:.06f}\")\n            del tr_ds, te_ds\n            gc.collect()\n            models[1] = model \n        if 'cat' in model_names:\n            print('cat model training started. It takes some minutes. Have a coffee!')\n            tr_ds = cat.Pool(tr_X, label=tr_y, weight=tr_weight)\n            te_ds = cat.Pool(val_X, label=val_y, weight=val_weight)\n            model = cat.CatBoostRegressor(**CAT_PARAMS)\n            model.fit(tr_ds, eval_set=te_ds, verbose=False)\n            print(f\"cat: {model.best_score_['validation']['r2_cbt']:.06f}\")\n            del tr_ds, te_ds\n            gc.collect()\n            models[2] = model\n\n        return models\n\n    def infer(data, models):\n        return np.mean([model.predict(data) for model in models], axis=0)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.808063Z","iopub.execute_input":"2024-10-22T11:14:51.808352Z","iopub.status.idle":"2024-10-22T11:14:51.824759Z","shell.execute_reply.started":"2024-10-22T11:14:51.808321Z","shell.execute_reply":"2024-10-22T11:14:51.823926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:\n\n    # df = get_df()\n    # models = train_model(df, MODEL_NAMES)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.825724Z","iopub.execute_input":"2024-10-22T11:14:51.826019Z","iopub.status.idle":"2024-10-22T11:14:51.837976Z","shell.execute_reply.started":"2024-10-22T11:14:51.825989Z","shell.execute_reply":"2024-10-22T11:14:51.837138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:    \n    \n    count = 0\n    models: list = None\n    lags_ : pl.DataFrame | None = None","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.839069Z","iopub.execute_input":"2024-10-22T11:14:51.840971Z","iopub.status.idle":"2024-10-22T11:14:51.855023Z","shell.execute_reply.started":"2024-10-22T11:14:51.840929Z","shell.execute_reply":"2024-10-22T11:14:51.854100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_3' in ENSEMBLE_SOLUTIONS:\n    \n    # Replace this function with your inference code.\n    # You can return either a Pandas or Polars dataframe, though Polars is recommended.\n    # Each batch of predictions (except the very first) must be returned within 10 minutes of the batch features being provided.\n    def predict_3(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n        \"\"\"Make a prediction.\"\"\"\n        # All the responders from the previous day are passed in at time_id == 0. We save them in a global variable for access at every time_id.\n        # Use them as extra features, if you like.\n        global lags_, models, count \n        if lags is not None:\n            lags_ = lags\n\n        if count == 0:\n            print('[1] Loading data...')\n            df = get_df()\n            print('data-shape:', df.shape)\n            print('[2] Training started')\n            models = train_model(df, MODEL_NAMES)\n            del df\n            gc.collect()\n\n        count += 1\n\n        predictions = test.select('row_id',pl.lit(0.0).alias('responder_6'))\n\n        feat = test[FEAT_COLS].to_pandas()\n        lgb_pred = models[0].predict(feat)\n        # xgb_pred = models[1].predict(xgb.DMatrix(feat))\n        # cat_pred = models[2].predict(feat)\n        # pred = [lgb_pred, xgb_pred, cat_pred]\n        # pred = np.mean(pred, axis=0)\n        pred = lgb_pred\n\n        predictions = predictions.with_columns(pl.Series('responder_6', pred))\n\n        # The predict function must return a DataFrame\n        assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n        # with columns 'row_id', 'responer_6'\n        assert predictions.columns == ['row_id', 'responder_6']\n        # and as many rows as the test data.\n        assert len(predictions) == len(test)\n\n        return predictions","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.856243Z","iopub.execute_input":"2024-10-22T11:14:51.858254Z","iopub.status.idle":"2024-10-22T11:14:51.866913Z","shell.execute_reply.started":"2024-10-22T11:14:51.858207Z","shell.execute_reply":"2024-10-22T11:14:51.866109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. [Jane Street:Cat_base](https://www.kaggle.com/code/backpaker/jane-street-cat-base) Lb=0.0023\n### [Backpacker](https://www.kaggle.com/backpaker)","metadata":{}},{"cell_type":"code","source":"if 'SOLUTION_4' in ENSEMBLE_SOLUTIONS:\n    \n    import os\n    import joblib \n\n    import pandas as pd\n    import polars as pl\n    import lightgbm as lgb\n    import xgboost as xgb\n    from catboost import CatBoostRegressor, Pool\n    import numpy as np \n\n    from joblib import Parallel, delayed\n\n    import kaggle_evaluation.jane_street_inference_server\n    \n    \n    test = pl.read_parquet(\"/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet/date_id=0/part-0.parquet\")\n    \n    cat_models = joblib.load(\"/kaggle/input/jscat-baseline/cat_models_baseline.pkl\")\n    \n    display(test)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:14:51.868003Z","iopub.execute_input":"2024-10-22T11:14:51.868376Z","iopub.status.idle":"2024-10-22T11:15:57.246768Z","shell.execute_reply.started":"2024-10-22T11:14:51.868345Z","shell.execute_reply":"2024-10-22T11:15:57.245845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_4' in ENSEMBLE_SOLUTIONS:\n\n    def predict_4(test: pl.DataFrame, lags: pl.DataFrame | None) -> pl.DataFrame | pd.DataFrame:\n\n        preds = np.zeros(test.shape[0])\n\n        for model in cat_models:\n            pred = model.predict(test.to_pandas())\n            preds += pred\n\n        mean_pred = preds / len(cat_models)\n\n        predictions = test.select(\n            'row_id',\n            pl.lit(0.0).alias('responder_6'),\n        )\n\n        predictions = predictions.with_columns(pl.Series('responder_6', mean_pred.ravel()))\n\n        return predictions","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:15:57.248198Z","iopub.execute_input":"2024-10-22T11:15:57.248542Z","iopub.status.idle":"2024-10-22T11:15:57.255191Z","shell.execute_reply.started":"2024-10-22T11:15:57.248508Z","shell.execute_reply":"2024-10-22T11:15:57.254322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## [JS Ridge baseline](https://www.kaggle.com/code/yunsuxiaozi/js-ridge-baseline) Lb=0.0026\n### [yunsuxiaozi](https://www.kaggle.com/yunsuxiaozi)","metadata":{}},{"cell_type":"markdown","source":"#### Created by <a href=\"https://github.com/yunsuxiaozi/\">github.com/yunsuxiaozi </a>  2024/10/16\n\n#### We will use Ridge as the baseline here.","metadata":{"_kg_hide-input":true}},{"cell_type":"code","source":"if 'SOLUTION_5' in ENSEMBLE_SOLUTIONS:\n\n    #necessary\n    import polars as pl                    # 和pandas类似,但是处理大型数据集有更好的性能.\n    import pandas as pd                    # 导入csv文件的库\n    import numpy as np                     # 对矩阵进行科学计算的库\n    \n    #model\n    from sklearn.linear_model import Ridge # 岭回归模型\n    import os                              # 与操作系统进行交互的库\n    import warnings                        # 避免一些可以忽略的报错\n    \n    warnings.filterwarnings('ignore')      # filterwarnings() 方法是用于设置警告过滤器的方法，它可以控制警告信息的输出方式和级别。\n           \n    import kaggle_evaluation.jane_street_inference_server  # 比赛方提供的环境\n\n    import random                          # 提供了一些用于生成随机数的函数\n                                           \n    def seed_everything(seed):             # 设置随机种子,保证模型可以复现\n        np.random.seed(seed)               # numpy的随机种子\n        random.seed(seed)                  # python内置的随机种子\n        \n    seed_everything(seed=2024)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-10-22T11:15:57.256568Z","iopub.execute_input":"2024-10-22T11:15:57.256855Z","iopub.status.idle":"2024-10-22T11:15:57.270815Z","shell.execute_reply.started":"2024-10-22T11:15:57.256823Z","shell.execute_reply":"2024-10-22T11:15:57.269837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_5' in ENSEMBLE_SOLUTIONS:\n    \n    def custom_metric(y_true,y_pred,weight):\n        weighted_r2=1-(np.sum(weight*(y_true-y_pred)**2)/np.sum(weight*y_true**2))\n        return weighted_r2\n\n    print(\"read data\")\n    train=pl.read_parquet(\"/kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet/partition_id=9/part-0.parquet\")\n    train=train.to_pandas()\n    print(\"get X,y\")\n\n    cols=[f'feature_0{i}' if i<10 else f'feature_{i}' for i in range(79)]\n    X=train[cols].fillna(3).values\n    y=train['responder_6'].values\n    print(\"train test split\")\n    split=1300000#大约是8:2\n    weights=train['weight'].values\n    train_X,train_y,test_X,test_y,train_weight,test_weight=X[:-split],y[:-split],X[-split:],y[-split:],weights[:-split],weights[-split:]\n    print(f\"train_X.shape:{train_X.shape},test_X.shape:{test_X.shape}\")\n    print(\"fit and predict\")\n    model=Ridge()\n    model.fit(train_X,train_y)\n    train_pred=model.predict(train_X)\n    test_pred=model.predict(test_X)\n    print(f\"train weighted_r2:{custom_metric(train_y,train_pred,weight=train_weight)}\")\n    print(f\"test weighted_r2:{custom_metric(test_y,test_pred,weight=test_weight)}\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-10-22T11:15:57.275617Z","iopub.execute_input":"2024-10-22T11:15:57.275906Z","iopub.status.idle":"2024-10-22T11:16:07.287232Z","shell.execute_reply.started":"2024-10-22T11:15:57.275875Z","shell.execute_reply":"2024-10-22T11:16:07.285972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if 'SOLUTION_5' in ENSEMBLE_SOLUTIONS:\n    \n    def predict_5(test,lags):\n        cols=[f'feature_0{i}' if i<10 else f'feature_{i}' for i in range(79)]\n        predictions = test.select(\n            'row_id',\n            pl.lit(0.0).alias('responder_6'),\n        )\n        test_preds=model.predict(test[cols].to_pandas().fillna(3).values)\n        predictions = predictions.with_columns(pl.Series('responder_6', test_preds.ravel()))\n        return predictions","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-10-22T11:16:07.289326Z","iopub.execute_input":"2024-10-22T11:16:07.290190Z","iopub.status.idle":"2024-10-22T11:16:07.300181Z","shell.execute_reply.started":"2024-10-22T11:16:07.290123Z","shell.execute_reply":"2024-10-22T11:16:07.298961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"####  We can replace Ridge with a more complex neural network.\n2024, Copyright [yunsuxiaozi]((https://www.kaggle.com/yunsuxiaozi))","metadata":{"_kg_hide-input":true}},{"cell_type":"markdown","source":"## Ensemble of solution","metadata":{}},{"cell_type":"code","source":"def predict(\n    test: pl.DataFrame, \n    lags: pl.DataFrame | None\n) -> pl.DataFrame | pd.DataFrame:\n    \"This is the inference and submission function used to predict the test set for the competition\"\n    \n    df_model_1 = predict_1(test,lags)\n    df_model_2 = predict_2(test,lags)\n    df_model_3 = predict_3(test,lags)\n    df_model_4 = predict_4(test,lags)\n    df_model_5 = predict_5(test,lags)\n    \n    # List of dataframes\n    df_list = [df_model_1, df_model_2, df_model_3, df_model_4, df_model_5]\n\n    # Accuracies of the models\n    accuracies = np.array([0.0044, 0.0040, 0.0034, 0.0023, 0.0026])\n\n    # Normalize the accuracies to get weights\n    weights = accuracies / accuracies.sum()\n\n    # Merge the dataframes on 'row_id'\n    merged_df = df_model_1[['row_id']].copy()\n    for i, df in enumerate(df_list):\n        merged_df = merged_df.merge(df[['row_id', 'responder_6']], on='row_id', suffixes=('', f'_model_{i+1}'))\n\n    # Function to compute the weighted vote\n    def weighted_average(row, weights):\n        predictions = row[1:].values  # Exclude 'row_id' from the row\n        weighted_preds = np.dot(weights, predictions)\n        return weighted_preds\n\n\n    # Apply the weighted vote row by row\n    merged_df['ensemble_prediction'] = merged_df.apply(lambda row: weighted_vote(row, weights), axis=1)\n    \n    df.rename(columns={'ensemble_prediction': 'responder_6'}, inplace=True)\n\n    # Final result\n    return merged_df[['row_id', 'responder_6']]\n    \n#     import random\n    \n#     if random.choice(ENSEMBLE_SOLUTIONS) == 'SOLUTION_2': return predict_2(test,lags)\n#     if random.choice(ENSEMBLE_SOLUTIONS) == 'SOLUTION_3': return predict_3(test,lags)\n#     if random.choice(ENSEMBLE_SOLUTIONS) == 'SOLUTION_4': return predict_4(test,lags)\n#     if random.choice(ENSEMBLE_SOLUTIONS) == 'SOLUTION_5': return predict_5(test,lags)","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2024-10-22T11:16:07.302876Z","iopub.execute_input":"2024-10-22T11:16:07.303875Z","iopub.status.idle":"2024-10-22T11:16:07.321702Z","shell.execute_reply.started":"2024-10-22T11:16:07.303809Z","shell.execute_reply":"2024-10-22T11:16:07.320458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inference_server = kaggle_evaluation.jane_street_inference_server.JSInferenceServer(predict)\n\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n    inference_server.serve()\nelse:\n    inference_server.run_local_gateway(\n        (\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/test.parquet',\n            '/kaggle/input/jane-street-real-time-market-data-forecasting/lags.parquet',\n        )\n    )","metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2024-10-22T11:16:07.323958Z","iopub.execute_input":"2024-10-22T11:16:07.325039Z","iopub.status.idle":"2024-10-22T11:16:07.480768Z","shell.execute_reply.started":"2024-10-22T11:16:07.324972Z","shell.execute_reply":"2024-10-22T11:16:07.479523Z"},"trusted":true},"execution_count":null,"outputs":[]}]}