{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":105399,"databundleVersionId":12733338,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# FlightRank XGB Ranker: Explanation","metadata":{}},{"cell_type":"markdown","source":"https://www.kaggle.com/code/antonoof/baseline-xgboost","metadata":{}},{"cell_type":"markdown","source":"\n\n### 🔷 **XGBRanker (XGBoost Ranker)**\n\n**Key Features:**\n\n* Implements **pairwise ranking** using the `rank:pairwise` objective.\n* Based on **gradient boosting over decision trees** (GBDT).\n* Provides **robust regularization** (`reg_alpha`, `reg_lambda`) to reduce overfitting.\n* Highly optimized for speed and performance (uses histogram-based algorithm).\n* Native support for **missing values** and **sparse inputs**.\n* GPU acceleration is available.\n\n**Pros:**\n\n* High performance in competitions and benchmarks.\n* Fine control over regularization and boosting behavior.\n* Well-documented and widely used in industry and academia.\n\n**Cons:**\n\n* Requires **manual preprocessing** for categorical features (e.g., one-hot encoding).\n* Slower than LightGBM for very large datasets.\n* Training can become unstable without proper group formatting.\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport polars as pl # read train -> pd\nimport xgboost as xgb\nfrom sklearn.model_selection import train_test_split","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:02:29.156052Z","iopub.execute_input":"2025-06-24T12:02:29.15632Z","iopub.status.idle":"2025-06-24T12:02:30.805594Z","shell.execute_reply.started":"2025-06-24T12:02:29.156287Z","shell.execute_reply":"2025-06-24T12:02:30.804906Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"features = ['bySelf', 'companyID', \n            'frequentFlyer', \n            'nationality', 'isAccess3D', 'isVip',\n            'legs0_segments0_baggageAllowance_quantity', \n            'legs0_segments0_baggageAllowance_weightMeasurementType', \n            'legs0_segments0_cabinClass',\n            'legs0_segments0_flightNumber', \n            'legs0_segments0_seatsAvailable', \n            'profileId', 'pricingInfo_isAccessTP', \n            'pricingInfo_passengerCount',\n            'legs0_segments0_arrivalTo_airport_iata', \n            'legs0_segments0_marketingCarrier_code',\n            'ranker_id', 'taxes', 'totalPrice'\n]\nfeatures_train = features + ['selected']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:02:30.806243Z","iopub.execute_input":"2025-06-24T12:02:30.806538Z","iopub.status.idle":"2025-06-24T12:02:30.81051Z","shell.execute_reply.started":"2025-06-24T12:02:30.806511Z","shell.execute_reply":"2025-06-24T12:02:30.809864Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train = pl.read_parquet('/kaggle/input/aeroclub-recsys-2025/train.parquet')\ntrain = train.select(features_train)\ntrain = train.to_pandas()\ntest = pd.read_parquet('/kaggle/input/aeroclub-recsys-2025/test.parquet')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:02:30.812281Z","iopub.execute_input":"2025-06-24T12:02:30.812475Z","iopub.status.idle":"2025-06-24T12:03:05.890203Z","shell.execute_reply.started":"2025-06-24T12:02:30.81246Z","shell.execute_reply":"2025-06-24T12:03:05.88945Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n1. A list of column names is defined. These columns contain categorical information such as flight number, airport codes, and airline codes.\n\n2. Each of these columns in both the training and test datasets is converted to the \"category\" data type. This step helps reduce memory usage and prepares the data for encoding.\n\n3. The categorical columns are then encoded into integer codes. Each unique category is replaced with a unique integer in both datasets. This is necessary because many machine learning models only accept numerical input.\n","metadata":{}},{"cell_type":"code","source":"categorical = ['frequentFlyer', \n               'legs0_segments0_flightNumber', \n               'legs0_segments0_arrivalTo_airport_iata', \n               'legs0_segments0_marketingCarrier_code']\nfor col in categorical:\n    train[col] = train[col].astype('category')\n    test[col] = test[col].astype('category')\n\nfor col in categorical:\n    train[col] = train[col].cat.codes\n    test[col] = test[col].cat.codes","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:03:05.891109Z","iopub.execute_input":"2025-06-24T12:03:05.891357Z","iopub.status.idle":"2025-06-24T12:03:11.831268Z","shell.execute_reply.started":"2025-06-24T12:03:05.891338Z","shell.execute_reply":"2025-06-24T12:03:11.830621Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"1. The `ranker_id` column in the training dataset is converted to the \"category\" data type. This is useful for memory efficiency and for models that can handle categorical variables.\n\n2. The same conversion is applied to the `ranker_id` column in the test dataset.\n\n3. A new variable `X` is created, which contains only the feature columns from the training data.\n\n4. A new variable `y` is created, which holds the target labels (`selected`) from the training data. This is typically the variable the model will try to predict.\n","metadata":{}},{"cell_type":"code","source":"train['ranker_id'] = train['ranker_id'].astype('category')\ntest['ranker_id'] = test['ranker_id'].astype('category')\n\nX = train[features]\ny = train['selected']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:03:11.831985Z","iopub.execute_input":"2025-06-24T12:03:11.83221Z","iopub.status.idle":"2025-06-24T12:03:15.489818Z","shell.execute_reply.started":"2025-06-24T12:03:11.832192Z","shell.execute_reply":"2025-06-24T12:03:15.488911Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n1. **Save the original categories of `ranker_id`**:\n   The unique category labels from the full training data are stored, so they can be reused consistently during encoding.\n\n2. **Split the data**:\n   The feature matrix `X` and target `y` are split into training and testing sets using an 80/20 ratio.\n\n3. **Encode `ranker_id` in the training set**:\n   The categorical column `ranker_id` in the training split is converted into numerical codes.\n\n4. **Encode `ranker_id` in the test set using the same categories**:\n   The test split's `ranker_id` is set to use the same category mapping as the training set, ensuring consistency, and then converted to numerical codes.\n\n5. **Create group sizes for ranking models**:\n   The number of samples (rows) for each `ranker_id` in the training set is counted and sorted by the ranker ID. This list of counts will be used to define groups in ranking models like LightGBM Ranker.\n","metadata":{}},{"cell_type":"code","source":"train_categories = X['ranker_id'].cat.categories\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)\nX_train['ranker_id'] = X_train['ranker_id'].cat.codes\nX_test['ranker_id'] = X_test['ranker_id'].cat.set_categories(train_categories, ordered=True).cat.codes\ntrain_group_sizes = X_train['ranker_id'].value_counts().sort_index().tolist()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:03:15.490771Z","iopub.execute_input":"2025-06-24T12:03:15.491061Z","iopub.status.idle":"2025-06-24T12:03:21.729453Z","shell.execute_reply.started":"2025-06-24T12:03:15.491034Z","shell.execute_reply":"2025-06-24T12:03:21.728847Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n1. **Model definition**:\n   An XGBoost ranking model (`XGBRanker`) is created with the following settings:\n\n   * `objective='rank:pairwise'`: Optimizes pairwise ranking (used to order items correctly within a group).\n   * `n_estimators=3000`: The number of boosting rounds (trees).\n   * `max_depth=5`: The maximum depth of each tree.\n   * `learning_rate=0.002`: A small step size for updating weights—ensures slow, careful learning.\n   * `eval_metric='ndcg@3'`: Evaluation uses NDCG at rank 3 (measures the quality of the top 3 ranked results).\n   * `reg_alpha` and `reg_lambda`: Regularization parameters to prevent overfitting (L1 and L2 penalties).\n\n2. **Model training**:\n   The model is trained on the training data (`X_train`, `y_train`), and the `group` parameter is passed to define how the samples are grouped for ranking. Each group corresponds to a list of items (e.g., recommendations for a user) that should be ranked internally.\n","metadata":{}},{"cell_type":"code","source":"model = xgb.XGBRanker(\n    objective='rank:pairwise',\n    n_estimators=3000,\n    max_depth=5,\n    learning_rate=0.002,\n    eval_metric='ndcg@3',\n    reg_alpha=0.6,\n    reg_lambda=0.8\n)\n\nmodel.fit(X_train, y_train, group=train_group_sizes)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:03:21.730269Z","iopub.execute_input":"2025-06-24T12:03:21.730535Z","iopub.status.idle":"2025-06-24T12:39:03.585417Z","shell.execute_reply.started":"2025-06-24T12:03:21.730512Z","shell.execute_reply":"2025-06-24T12:39:03.584686Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n### **Function Purpose**\n\n`take_rank` assigns a rank to each item within a group based on a score, with higher scores ranked higher.\n\n---\n\n### **Parameters**\n\n* `data`: A DataFrame containing the data.\n* `group_col`: The column name used to define groups (e.g., user ID or ranker ID).\n* `score_col`: The column containing the score used for ranking.\n* `rank_col`: The name of the new column that will store the ranking result.\n\n---\n\n### **Inside the Function**\n\n* The DataFrame is grouped by `group_col`.\n* Within each group, the `score_col` values are ranked in descending order (higher scores get lower rank numbers).\n* `method='first'` means that if multiple items have the same score, their original order is preserved.\n* The resulting ranks are converted to integers and stored in a new column (`rank_col`).\n* Finally, the updated DataFrame is returned.\n\n---\n\n### **Example Use Case**\n\nYou might use this after model prediction to rank recommendations for each user based on predicted scores.\n","metadata":{}},{"cell_type":"code","source":"def take_rank(data, group_col, score_col, rank_col):\n    data[rank_col] = data.groupby(group_col, observed=False)[score_col].rank(method='first', ascending=False).astype(int)\n    return data","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:39:03.586202Z","iopub.execute_input":"2025-06-24T12:39:03.586737Z","iopub.status.idle":"2025-06-24T12:39:03.590342Z","shell.execute_reply.started":"2025-06-24T12:39:03.586711Z","shell.execute_reply":"2025-06-24T12:39:03.589758Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n---\n\n### **Function: `HitRate3`**\n\n**Goal**:\nTo calculate **HitRate\\@3**, which measures how often the true positive item is ranked in the top 3 predictions per group.\n\n---\n\n**Step-by-step inside `HitRate3`**:\n\n1. `positive = 0`: Counter for correct predictions.\n2. `Q`: Number of unique groups (e.g., number of users or queries).\n3. For each group (by `ranker_id`):\n\n   * Select the **top 3 items** with the **lowest rank values** (i.e., highest predicted scores).\n   * If any of these top 3 items has `true_col == 1` (i.e., was actually selected), count it as a **hit**.\n4. Finally, divide the number of hits by the total number of groups to get the **HitRate\\@3** score.\n\n---\n\n### **Evaluation Process**\n\n1. **Predict scores** for the test set using the trained model.\n2. **Rank the predictions** within each group using the predicted scores.\n3. **Assign ground-truth labels** (`selected`) from `y_test` to the ranked test data.\n4. **Compute the hit rate** using the `HitRate3` function.\n5. **Print the result** formatted as a percentage.\n","metadata":{}},{"cell_type":"code","source":"def HitRate3(data, rank_col, true_col):\n    positive = 0\n    Q = data['ranker_id'].nunique()\n\n    for _, group in data.groupby('ranker_id'):\n        preds_top_3 = group.nsmallest(3, rank_col) # top 3\n        if any(preds_top_3[true_col] == 1):\n            positive += 1\n\n    hit_rate = positive / Q\n    return hit_rate\n\nX_test['y_scores'] = model.predict(X_test[features])\nX_test = take_rank(X_test, 'ranker_id', 'y_scores', 'y_ranks')\nX_test['selected'] = y_test.values\n\nhitrate_score = HitRate3(X_test, 'y_ranks', 'selected')\nprint(f\"HitRate@3: {hitrate_score:.2f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:39:03.592019Z","iopub.execute_input":"2025-06-24T12:39:03.592199Z","iopub.status.idle":"2025-06-24T12:41:28.62647Z","shell.execute_reply.started":"2025-06-24T12:39:03.592186Z","shell.execute_reply":"2025-06-24T12:41:28.625617Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n1. **Predict scores on the test set**:\n   The trained model is used to generate predicted scores (`y_scores`) for each item in the test data.\n\n2. **Rank items within each group**:\n   The `take_rank` function ranks the items per `ranker_id` group based on these predicted scores. The ranks are stored in the `selected` column.\n\n3. **Prepare the submission file**:\n   A new DataFrame `submission` is created containing only the necessary columns: `Id`, `ranker_id`, and the computed `selected` rank (which will typically be used to determine the top choice).\n\n4. **Save as a Parquet file**:\n   The `submission` DataFrame is saved in Parquet format without including the index column.\n\n5. **Display the first few rows**:\n   `submission.head()` is called to visually check the top of the submission file.\n\n","metadata":{}},{"cell_type":"code","source":"test['y_scores'] = model.predict(test[features])\ntest = take_rank(test, 'ranker_id', 'y_scores', 'selected')\nsubmission = test[['Id', 'ranker_id', 'selected']]\nsubmission.to_parquet('submission.parquet', index=False)\nsubmission.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T12:41:28.627282Z","iopub.execute_input":"2025-06-24T12:41:28.62768Z","iopub.status.idle":"2025-06-24T12:42:27.253336Z","shell.execute_reply.started":"2025-06-24T12:41:28.627657Z","shell.execute_reply":"2025-06-24T12:42:27.2527Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}