{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.18","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"tpu1vmV38","dataSources":[{"sourceId":105399,"databundleVersionId":12733338,"sourceType":"competition"}],"dockerImageVersionId":31091,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# FlightRank CatBoost Ranker: Explanation\nhttps://www.kaggle.com/code/ka1242/catboost-ranker-baseline-flightrank-2025","metadata":{}},{"cell_type":"markdown","source":"**CatBoost Ranker** is a gradient boosting model developed by Yandex, specialized for ranking tasks in information retrieval or recommendation systems. It extends CatBoost’s gradient boosting framework to **ranking** by optimizing for **relative ordering of items** rather than direct predictions.\n\n### Key Concepts:\n\n* **Ranking Objective**: Unlike classification or regression, ranking focuses on ordering items—for example, showing the most relevant documents or products first. CatBoost Ranker supports ranking objectives like `YetiRank` and `QuerySoftMax`, which are designed for pairwise or listwise optimization.\n\n* **Group-Based Learning**: It requires grouping information—such as user sessions, search queries, or any context—in which ranking should be learned. This is typically passed as a group ID per instance.\n\n* **Handling Categorical Features**: Like all CatBoost models, it natively supports categorical features without preprocessing, which is a strong advantage over many other ranking models.\n\n* **Efficient Training**: CatBoost uses ordered boosting to avoid overfitting and improve generalization, particularly helpful in ranking tasks with noisy or sparse labels.\n\n### Use Cases:\n\n* Search engine result ranking\n* Recommendation systems (e.g., movies, products)\n* Personalized content delivery\n\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here's a **brief comparison of CatBoost Ranker vs LightGBM Ranker vs XGBoost Ranker** across key dimensions, assuming all are used for **learning-to-rank (LTR)** tasks.\n\n---\n\n### 🧠 **Core Algorithm**\n\n| Feature            | **CatBoost**                         | **LightGBM**                | **XGBoost**                              |\n| ------------------ | ------------------------------------ | --------------------------- | ---------------------------------------- |\n| Boosting           | Oblivious (symmetric) decision trees | Leaf-wise                   | Level-wise                               |\n| Ranking Objectives | `YetiRank`, `QuerySoftMax`           | `lambdarank`, `rank_xendcg` | `rank:pairwise`, `rank:ndcg`, `rank:map` |\n| Ranking Style      | Listwise or pairwise                 | Pairwise/Listwise           | Pairwise/Listwise                        |\n\n---\n\n### ⚙️ **Ease of Use**\n\n| Feature                    | **CatBoost** | **LightGBM**      | **XGBoost**       |\n| -------------------------- | ------------ | ----------------- | ----------------- |\n| Categorical Support        | Native       | Requires encoding | Requires encoding |\n| Handling of Missing Values | Automatic    | Automatic         | Automatic         |\n| GPU Support                | Yes          | Yes               | Yes               |\n\n---\n\n### 🏎️ **Speed & Efficiency**\n\n| Metric         | **CatBoost**                     | **LightGBM**                      | **XGBoost** |\n| -------------- | -------------------------------- | --------------------------------- | ----------- |\n| Training Speed | Slower (due to ordered boosting) | Fastest (leaf-wise is aggressive) | Medium      |\n| Memory Usage   | Moderate                         | Low                               | Moderate    |\n\n---\n\n### 📈 **Performance**\n\n| Feature            | **CatBoost**                          | **LightGBM**                     | **XGBoost** |\n| ------------------ | ------------------------------------- | -------------------------------- | ----------- |\n| Overfitting Risk   | Lower (ordered boosting helps)        | Higher (leaf-wise can overfit)   | Moderate    |\n| Accuracy (general) | High on noisy/sparse/categorical data | High on large-scale numeric data | Competitive |\n\n---\n\n### 📦 **Best Use Cases**\n\n* **CatBoost**: Tabular data with many categorical features; small to medium-sized datasets; cold-start friendly.\n* **LightGBM**: Large-scale datasets with numerical features; speed-critical applications.\n* **XGBoost**: Versatile, stable, and widely used; good balance between performance and interpretability.\n\n---\n\n### 🔚 Summary\n\n* Use **CatBoost Ranker** when you have **many categorical variables** and want strong default performance.\n* Use **LightGBM Ranker** for **very large datasets** and **high-speed needs**, but be careful with overfitting.\n* Use **XGBoost Ranker** for a **well-balanced, reliable baseline** with mature community support.\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Import necessary libraries\nimport pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import GroupShuffleSplit\nfrom sklearn.metrics import log_loss\nimport matplotlib.pyplot as plt\n\n# Global parameters\nTRAIN_SAMPLE_FRAC = 0.20  # Sample 30% of data for faster iteration\nRANDOM_STATE = 42\nnp.random.seed(RANDOM_STATE)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:37:31.711383Z","iopub.execute_input":"2025-06-23T13:37:31.71165Z","iopub.status.idle":"2025-06-23T13:37:37.679822Z","shell.execute_reply.started":"2025-06-23T13:37:31.71162Z","shell.execute_reply":"2025-06-23T13:37:37.674865Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load parquet files\ntrain = pd.read_parquet('/kaggle/input/aeroclub-recsys-2025/train.parquet')\ntest = pd.read_parquet('/kaggle/input/aeroclub-recsys-2025/test.parquet')\n\nprint(f\"Train shape: {train.shape}, Test shape: {test.shape}\")\nprint(f\"Unique ranker_ids in train: {train['ranker_id'].nunique():,}\")\nprint(f\"Selected rate: {train['selected'].mean():.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:37:37.909187Z","iopub.execute_input":"2025-06-23T13:37:37.909461Z","iopub.status.idle":"2025-06-23T13:38:10.537304Z","shell.execute_reply.started":"2025-06-23T13:37:37.909438Z","shell.execute_reply":"2025-06-23T13:38:10.531617Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This code performs **group-level sampling** for training a ranking model. Instead of randomly picking individual data rows, it randomly selects entire groups (identified by `ranker_id`). This ensures that all items within a group stay together, which is essential for ranking tasks where the model learns the relative ordering **within** each group.\n\nIt's used when you want to train on a smaller portion of the data (`<100%`) but still preserve the structure required for meaningful learning-to-rank behavior.\n","metadata":{}},{"cell_type":"code","source":"# Sample by ranker_id to keep groups intact\nif TRAIN_SAMPLE_FRAC < 1.0:\n    unique_rankers = train['ranker_id'].unique()\n    n_sample = int(len(unique_rankers) * TRAIN_SAMPLE_FRAC)\n    sampled_rankers = np.random.RandomState(RANDOM_STATE).choice(\n        unique_rankers, size=n_sample, replace=False\n    )\n    train = train[train['ranker_id'].isin(sampled_rankers)]\n    print(f\"Sampled train to {len(train):,} rows ({train['ranker_id'].nunique():,} groups)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:38:16.373721Z","iopub.execute_input":"2025-06-23T13:38:16.374014Z","iopub.status.idle":"2025-06-23T13:38:24.87059Z","shell.execute_reply.started":"2025-06-23T13:38:16.373988Z","shell.execute_reply":"2025-06-23T13:38:24.864804Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Convert ranker_id to string for CatBoost\ntrain['ranker_id'] = train['ranker_id'].astype(str)\ntest['ranker_id'] = test['ranker_id'].astype(str)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:38:25.984899Z","iopub.execute_input":"2025-06-23T13:38:25.985179Z","iopub.status.idle":"2025-06-23T13:38:26.183714Z","shell.execute_reply.started":"2025-06-23T13:38:25.985154Z","shell.execute_reply":"2025-06-23T13:38:26.178319Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"cat_features = [\n    'nationality', 'searchRoute', 'corporateTariffCode',\n    # Leg 0 segments 0-1\n    'legs0_segments0_aircraft_code', 'legs0_segments0_arrivalTo_airport_city_iata',\n    'legs0_segments0_arrivalTo_airport_iata', 'legs0_segments0_departureFrom_airport_iata',\n    'legs0_segments0_marketingCarrier_code', 'legs0_segments0_operatingCarrier_code',\n    'legs0_segments0_flightNumber',\n    'legs0_segments1_aircraft_code', 'legs0_segments1_arrivalTo_airport_city_iata',\n    'legs0_segments1_arrivalTo_airport_iata', 'legs0_segments1_departureFrom_airport_iata',\n    'legs0_segments1_marketingCarrier_code', 'legs0_segments1_operatingCarrier_code',\n    'legs0_segments1_flightNumber',\n    # Leg 1 segments 0-1\n    'legs1_segments0_aircraft_code', 'legs1_segments0_arrivalTo_airport_city_iata',\n    'legs1_segments0_arrivalTo_airport_iata', 'legs1_segments0_departureFrom_airport_iata',\n    'legs1_segments0_marketingCarrier_code', 'legs1_segments0_operatingCarrier_code',\n    'legs1_segments0_flightNumber',\n    'legs1_segments1_aircraft_code', 'legs1_segments1_arrivalTo_airport_city_iata',\n    'legs1_segments1_arrivalTo_airport_iata', 'legs1_segments1_departureFrom_airport_iata',\n    'legs1_segments1_marketingCarrier_code', 'legs1_segments1_operatingCarrier_code',\n    'legs1_segments1_flightNumber'\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:38:27.180631Z","iopub.execute_input":"2025-06-23T13:38:27.18097Z","iopub.status.idle":"2025-06-23T13:38:27.193229Z","shell.execute_reply.started":"2025-06-23T13:38:27.180943Z","shell.execute_reply":"2025-06-23T13:38:27.187343Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This function `create_features(df)` performs extensive **feature engineering** to enrich a flight ranking dataset with meaningful variables that help models (like CatBoost or LightGBM rankers) better understand user preferences. \n\n---\n\n### 🔧 **Main Purpose**\n\nTo create new features from raw flight booking data that capture **price, duration, schedule, carrier, baggage, fees, and frequent flyer information**, while respecting the **ranking group structure** (`ranker_id`).\n\n---\n\n### 📊 **Key Feature Categories**\n\n1. **Duration Conversion**\n\n   * Converts time strings (like `\"01:45:00\"`) into total minutes.\n\n2. **Price-Based Features**\n\n   * Log-transformed prices, price-to-tax ratio, and relative price within a group.\n\n3. **Duration and Segment Features**\n\n   * Total trip duration, number of flight segments per leg, and one-way/return indicators.\n\n4. **Group Ranking Features**\n\n   * Rank of price and duration **within each `ranker_id` group**, indicators for cheapest/most expensive flights.\n\n5. **Frequent Flyer Info**\n\n   * Counts number of loyalty programs, matches with marketing carrier, and identifies VIP travelers.\n\n6. **Baggage and Fees**\n\n   * Presence and amount of baggage allowance, extra fees, and their proportion of total price.\n\n7. **Time-of-Day Features**\n\n   * Departure/arrival hour and weekday, as well as if the flight is during typical business hours.\n\n8. **Direct Flight Flags**\n\n   * Whether each leg is direct and if it’s the cheapest direct flight in the group.\n\n9. **Miscellaneous**\n\n   * Group size, popular routes, major carriers, cabin class info, and access-type flags.\n\n10. **Missing Value Handling**\n\n    * Fills missing numeric values with `0` and strings with `\"missing\"`.\n\n---\n\n### ✅ **Why It's Useful**\n\nThese engineered features give machine learning models much more **contextual understanding** of what makes one flight option more desirable than another—beyond just price or duration—while preserving **group integrity** essential for learning to rank.\n\n","metadata":{}},{"cell_type":"code","source":"def create_features(df):\n    \"\"\"\n    Return a copy of df enriched with engineered features for flight-ranking models.\n    \"\"\"\n    df = df.copy()\n\n    def hms_to_minutes(s: pd.Series) -> np.ndarray:\n        \"\"\"Vectorised 'HH:MM:SS' → minutes (seconds ignored).\"\"\"\n        mask = s.notna()\n        out = np.zeros(len(s), dtype=float)\n        if mask.any():\n            parts = s[mask].astype(str).str.split(':', expand=True)\n            out[mask] = (\n                pd.to_numeric(parts[0], errors=\"coerce\").fillna(0) * 60\n                + pd.to_numeric(parts[1], errors=\"coerce\").fillna(0)\n            )\n        return out\n\n    # Duration columns\n    dur_cols = (\n        [\"legs0_duration\", \"legs1_duration\"]\n        + [f\"legs{l}_segments{s}_duration\" for l in (0, 1) for s in (0, 1)]\n    )\n    for col in dur_cols:\n        if col in df.columns:\n            df[col] = hms_to_minutes(df[col])\n\n    # Feature container\n    feat = {}\n\n    # Price\n    feat[\"price_per_tax\"] = df[\"totalPrice\"] / (df[\"taxes\"] + 1)\n    feat[\"tax_rate\"] = df[\"taxes\"] / (df[\"totalPrice\"] + 1)\n    feat[\"log_price\"] = np.log1p(df[\"totalPrice\"])\n\n    # Durations\n    df[\"total_duration\"] = df[\"legs0_duration\"].fillna(0) + df[\"legs1_duration\"].fillna(0)\n    feat[\"duration_ratio\"] = np.where(\n        df[\"legs1_duration\"].fillna(0) > 0,\n        df[\"legs0_duration\"] / (df[\"legs1_duration\"] + 1),\n        1.0,\n    )\n\n    # Segment counts\n    for leg in (0, 1):\n        seg_cols = [f\"legs{leg}_segments{i}_duration\" for i in (0, 1)]\n        feat[f\"n_segments_leg{leg}\"] = df[seg_cols].notna().sum(axis=1)\n    feat[\"total_segments\"] = feat[\"n_segments_leg0\"] + feat[\"n_segments_leg1\"]\n\n    # Trip type\n    feat[\"is_one_way\"] = df[\"legs1_duration\"].isna().astype(int)\n\n    # Rank features\n    grp = df.groupby(\"ranker_id\")\n    feat[\"price_rank\"] = grp[\"totalPrice\"].rank()\n    feat[\"price_pct_rank\"] = grp[\"totalPrice\"].rank(pct=True)\n    feat[\"duration_rank\"] = grp[\"total_duration\"].rank()\n    feat[\"is_cheapest\"] = (grp[\"totalPrice\"].transform(\"min\") == df[\"totalPrice\"]).astype(int)\n    feat[\"is_most_expensive\"] = (grp[\"totalPrice\"].transform(\"max\") == df[\"totalPrice\"]).astype(int)\n    feat[\"price_from_median\"] = grp[\"totalPrice\"].transform(\n        lambda x: (x - x.median()) / (x.std() + 1)\n    )\n\n    # Frequent-flyer\n    ff = df[\"frequentFlyer\"].fillna(\"\").astype(str)\n    feat[\"n_ff_programs\"] = ff.str.count(\"/\") + (ff != \"\")\n    airlines = [\"SU\", \"S7\", \"U6\", \"TK\", \"DP\", \"UT\", \"EK\", \"N4\", \"5N\", \"LH\"]\n    for al in airlines:\n        feat[f\"ff_{al}\"] = ff.str.contains(rf\"\\b{al}\\b\").astype(int)\n    feat[\"ff_matches_carrier\"] = np.select(\n        [\n            (feat[f\"ff_{al}\"] == 1)\n            & (df[\"legs0_segments0_marketingCarrier_code\"] == al)\n            for al in [\"SU\", \"S7\", \"U6\", \"TK\"]\n        ],\n        [1, 1, 1, 1],\n        default=0,\n    )\n\n    # Binary flags\n    feat.update(\n        dict(\n            is_vip_freq=((df[\"isVip\"] == 1) | (feat[\"n_ff_programs\"] > 0)).astype(int),\n            has_return=(~df[\"legs1_duration\"].isna()).astype(int),\n            has_corporate_tariff=(~df[\"corporateTariffCode\"].isna()).astype(int),\n        )\n    )\n\n    # Baggage and fees\n    feat[\"baggage_total\"] = (\n        df[\"legs0_segments0_baggageAllowance_quantity\"].fillna(0)\n        + df[\"legs1_segments0_baggageAllowance_quantity\"].fillna(0)\n    )\n    feat[\"has_baggage\"] = (feat[\"baggage_total\"] > 0).astype(int)\n    feat[\"total_fees\"] = (\n        df[\"miniRules0_monetaryAmount\"].fillna(0) + df[\"miniRules1_monetaryAmount\"].fillna(0)\n    )\n    feat[\"has_fees\"] = (feat[\"total_fees\"] > 0).astype(int)\n    feat[\"fee_rate\"] = feat[\"total_fees\"] / (df[\"totalPrice\"] + 1)\n\n    # Time-of-day\n    for col in (\"legs0_departureAt\", \"legs0_arrivalAt\", \"legs1_departureAt\", \"legs1_arrivalAt\"):\n        if col in df.columns:\n            dt = pd.to_datetime(df[col], errors=\"coerce\")\n            feat[f\"{col}_hour\"] = dt.dt.hour.fillna(12)\n            feat[f\"{col}_weekday\"] = dt.dt.weekday.fillna(0)\n            h = dt.dt.hour.fillna(12)\n            feat[f\"{col}_business_time\"] = (((6 <= h) & (h <= 9)) | ((17 <= h) & (h <= 20))).astype(int)\n\n    # Direct-flight flags\n    feat[\"is_direct_leg0\"] = (feat[\"n_segments_leg0\"] == 1).astype(int)\n    feat[\"is_direct_leg1\"] = (feat[\"n_segments_leg1\"] == 1).astype(int)\n    feat[\"both_direct\"] = feat[\"is_direct_leg0\"] & feat[\"is_direct_leg1\"]\n\n    # Cheapest direct\n    df[\"_direct\"] = feat[\"n_segments_leg0\"] == 1\n    direct_min_price = df.loc[df[\"_direct\"]].groupby(\"ranker_id\")[\"totalPrice\"].min()\n    feat[\"is_direct_cheapest\"] = (\n        df[\"_direct\"] & (df[\"totalPrice\"] == df[\"ranker_id\"].map(direct_min_price))\n    ).astype(int)\n    df.drop(columns=\"_direct\", inplace=True)\n\n    # Misc flags\n    feat[\"has_access_tp\"] = (df[\"pricingInfo_isAccessTP\"] == 1).astype(int)\n    feat[\"group_size\"] = df.groupby(\"ranker_id\")[\"Id\"].transform(\"count\")\n    feat[\"group_size_log\"] = np.log1p(feat[\"group_size\"])\n    feat[\"is_major_carrier\"] = df[\"legs0_segments0_marketingCarrier_code\"].isin([\"SU\", \"S7\", \"U6\"]).astype(int)\n    popular_routes = {\"MOWLED/LEDMOW\", \"LEDMOW/MOWLED\", \"MOWLED\", \"LEDMOW\", \"MOWAER/AERMOW\"}\n    feat[\"is_popular_route\"] = df[\"searchRoute\"].isin(popular_routes).astype(int)\n    feat[\"avg_cabin_class\"] = df[[\"legs0_segments0_cabinClass\", \"legs1_segments0_cabinClass\"]].mean(axis=1)\n    feat[\"cabin_class_diff\"] = (\n        df[\"legs0_segments0_cabinClass\"].fillna(0) - df[\"legs1_segments0_cabinClass\"].fillna(0)\n    )\n\n    # Merge new features\n    df = pd.concat([df, pd.DataFrame(feat, index=df.index)], axis=1)\n\n    # Final NaN handling (loop avoids duplicate-column error)\n    for col in df.select_dtypes(include=\"number\").columns:\n        df[col] = df[col].fillna(0)\n    for col in df.select_dtypes(include=\"object\").columns:\n        df[col] = df[col].fillna(\"missing\")\n\n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:38:27.836866Z","iopub.execute_input":"2025-06-23T13:38:27.837154Z","iopub.status.idle":"2025-06-23T13:38:27.858722Z","shell.execute_reply.started":"2025-06-23T13:38:27.837129Z","shell.execute_reply":"2025-06-23T13:38:27.854432Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply feature engineering\ntrain = create_features(train)\ntest = create_features(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:38:29.554663Z","iopub.execute_input":"2025-06-23T13:38:29.554963Z","iopub.status.idle":"2025-06-23T13:40:42.761947Z","shell.execute_reply.started":"2025-06-23T13:38:29.554936Z","shell.execute_reply":"2025-06-23T13:40:42.757144Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Exclude columns\nexclude_cols = ['Id', 'ranker_id', 'selected', 'profileId', 'requestDate',\n                'legs0_departureAt', 'legs0_arrivalAt', 'legs1_departureAt', 'legs1_arrivalAt',\n                'miniRules0_percentage', 'miniRules1_percentage',  # >90% missing\n                'frequentFlyer']  # Already processed\n\n# Exclude segment 2-3 columns (>98% missing)\nfor leg in [0, 1]:\n    for seg in [2, 3]:\n        for suffix in ['aircraft_code', 'arrivalTo_airport_city_iata', 'arrivalTo_airport_iata',\n                      'baggageAllowance_quantity', 'baggageAllowance_weightMeasurementType',\n                      'cabinClass', 'departureFrom_airport_iata', 'duration', 'flightNumber',\n                      'marketingCarrier_code', 'operatingCarrier_code', 'seatsAvailable']:\n            exclude_cols.append(f'legs{leg}_segments{seg}_{suffix}')\n\nfeature_cols = [col for col in train.columns if col not in exclude_cols]\ncat_features_final = [col for col in cat_features if col in feature_cols]\n\nprint(f\"Using {len(feature_cols)} features ({len(cat_features_final)} categorical)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:41:00.824764Z","iopub.execute_input":"2025-06-23T13:41:00.825065Z","iopub.status.idle":"2025-06-23T13:41:00.83817Z","shell.execute_reply.started":"2025-06-23T13:41:00.82504Z","shell.execute_reply":"2025-06-23T13:41:00.832597Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This snippet prepares training, validation, and test data for a **learning-to-rank model** with **group-aware splitting**, which is crucial for maintaining the integrity of ranking tasks.\n\n---\n\n### 🔍 **Brief Explanation (No Code)**\n\n1. **Feature Selection**:\n   Selects the input features (`X`) and target labels (`y`, typically the relevance or rank) from the `train` dataset, along with group labels (`ranker_id`), which define the individual ranking tasks (e.g., flight searches).\n\n2. **Test Set**:\n   Prepares the test data (`X_test`) and its group structure, but without labels since it's for final prediction.\n\n3. **Group-Based Train/Validation Split**:\n   Uses `GroupShuffleSplit` to split the training data into **train and validation sets**, ensuring that **no `ranker_id` appears in both sets**. This prevents data leakage and simulates real-world generalization better.\n\n4. **Final Assignment**:\n   Extracts the actual row subsets for training and validation from the split indices.\n\n5. **Output Summary**:\n   Prints how many rows were allocated to training, validation, and test sets.\n\n---\n\n### ✅ Why It's Important\n\n* Ensures that all examples from a single group (`ranker_id`) are **only in one set**, preserving the integrity of ranking problems.\n* Prevents leakage of group-specific patterns into validation.\n* Helps ranking models like **CatBoost Ranker**, **LightGBM Ranker**, and **XGBoost Ranker** generalize correctly.\n\n","metadata":{}},{"cell_type":"code","source":"# Prepare data\nX_train = train[feature_cols]\ny_train = train['selected']\ngroups_train = train['ranker_id']\n\nX_test = test[feature_cols]\ngroups_test = test['ranker_id']\n\n# Group-based split\ngss = GroupShuffleSplit(n_splits=1, test_size=0.2, random_state=RANDOM_STATE)\ntrain_idx, val_idx = next(gss.split(X_train, y_train, groups_train))\n\nX_tr, X_val = X_train.iloc[train_idx], X_train.iloc[val_idx]\ny_tr, y_val = y_train.iloc[train_idx], y_train.iloc[val_idx]\ngroups_tr, groups_val = groups_train.iloc[train_idx], groups_train.iloc[val_idx]\n\nprint(f\"Train: {len(X_tr):,} rows, Val: {len(X_val):,} rows, Test: {len(X_test):,} rows\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:41:03.617318Z","iopub.execute_input":"2025-06-23T13:41:03.617811Z","iopub.status.idle":"2025-06-23T13:41:11.862457Z","shell.execute_reply.started":"2025-06-23T13:41:03.617752Z","shell.execute_reply":"2025-06-23T13:41:11.857824Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here’s a **brief explanation** of this code block that sets up and trains a **CatBoost Ranker** model for a learning-to-rank task:\n\n---\n\n### 🔧 **1. CatBoost Pools with Group Info**\n\n* Creates **`Pool` objects** for training and validation data, which is CatBoost’s optimized data format.\n* Includes:\n\n  * Feature matrix (`X`)\n  * Target labels (`y`)\n  * Group identifiers (`ranker_id`)\n  * List of categorical features (`cat_features_final`)\n* The **`group_id`** is essential for ranking tasks to indicate which rows belong to the same ranking query or context.\n\n---\n\n### 🤖 **2. CatBoostRanker Model Setup**\n\n* Initializes a `CatBoostRanker` model with the following key parameters:\n\n  * `loss_function='YetiRank'`: A listwise ranking loss designed for implicit feedback (like clicks).\n  * `eval_metric='PrecisionAt:top=3'`: Evaluates how often the top 3 predictions are relevant.\n  * `early_stopping_rounds=100`: Stops training if performance doesn't improve after 100 rounds.\n  * `cat_features`: Tells CatBoost which features are categorical—no need for one-hot encoding.\n  * `task_type='CPU'`: Can be changed to `'GPU'` if available for faster training.\n\n---\n\n### 📈 **3. Model Training**\n\n* Trains the model using `.fit()` on the training pool, with validation monitored via `eval_set`.\n* Uses `use_best_model=True` to retain the best iteration found on the validation set.\n\n---\n\n### ✅ **Why It’s Good Practice**\n\n* Respects ranking group structure via `group_id`.\n* Leverages CatBoost’s native handling of categorical features.\n* Uses early stopping and ranking-specific evaluation for robust performance.\n","metadata":{}},{"cell_type":"code","source":"!pip install -U catboost","metadata":{"trusted":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from catboost import CatBoostRanker, Pool","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create CatBoost pools\ntrain_pool = Pool(X_tr, y_tr, group_id=groups_tr, cat_features=cat_features_final)\nval_pool = Pool(X_val, y_val, group_id=groups_val, cat_features=cat_features_final)\n\n# Initialize CatBoost Ranker\nmodel = CatBoostRanker(\n    loss_function='YetiRank',\n    iterations=1000,\n    learning_rate=0.02,\n    depth=8,\n    l2_leaf_reg=0.2,\n    random_seed=RANDOM_STATE,\n    eval_metric='PrecisionAt:top=3',\n    early_stopping_rounds=100,\n    verbose=20,\n    task_type='CPU',  # Change to 'GPU' if available\n    cat_features=cat_features_final,\n#     grow_policy='Lossguide'\n)\n\n# Train model\nmodel.fit(train_pool, eval_set=val_pool, use_best_model=True);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:42:36.546498Z","iopub.execute_input":"2025-06-23T13:42:36.546884Z","iopub.status.idle":"2025-06-23T13:42:46.662032Z","shell.execute_reply.started":"2025-06-23T13:42:36.546854Z","shell.execute_reply":"2025-06-23T13:42:46.657387Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This code performs **post-processing and evaluation** for a CatBoost ranking model's predictions. Here's a concise, non-code explanation:\n\n---\n\n### 🧠 **1. Sigmoid Conversion**\n\n* The **raw prediction scores** from CatBoost Ranker are not probabilities.\n* A **scaled sigmoid function** (dividing by 10) is applied to smooth and squash these scores into the \\[0, 1] range.\n* This helps when you want to **normalize scores** for interpretation, ranking comparison, or thresholding.\n\n---\n\n### 🎯 **2. HitRate\\@K Evaluation**\n\n* **HitRate\\@k** checks if the **correct (selected)** item appears among the model’s **top-k predictions** within each group (`ranker_id`).\n* It evaluates **how often the model ranks a relevant item within the top-k slots**, which is crucial in recommendation systems or search ranking.\n* The function:\n\n  * Only includes groups with more than 10 options (to avoid trivial groups).\n  * Computes the fraction of such groups where the selected (ground-truth) item is ranked in the top-**k**.\n  * A higher value means the model ranks relevant items better.\n\n---\n\n### ✅ **Why It Matters**\n\n* Normalizing scores with a sigmoid helps with **comparability across groups**.\n* HitRate\\@k is an **intuitive and practical ranking metric**, reflecting **user-facing performance** (e.g., was the correct flight in the top 3?).\n\n","metadata":{}},{"cell_type":"code","source":"# Convert scores to probabilities using sigmoid\ndef sigmoid(x):\n    return 1 / (1 + np.exp(-x / 10))  # Scale factor for CatBoost scores\n\n# HitRate@3 calculation\ndef calculate_hitrate_at_k(df, k=3):\n    \"\"\"Calculate HitRate@k for groups with >10 options\"\"\"\n    hits = []\n    for ranker_id, group in df.groupby('ranker_id'):\n        # Only consider groups with >10 options\n        if len(group) > 10:\n            # Get top-k predictions\n            top_k = group.nlargest(k, 'pred')\n            # Check if selected item is in top-k\n            hit = (top_k['selected'] == 1).any()\n            hits.append(hit)\n    return np.mean(hits) if hits else 0.0","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-24T01:10:22.863802Z","iopub.execute_input":"2025-06-24T01:10:22.864172Z","iopub.status.idle":"2025-06-24T01:10:22.878514Z","shell.execute_reply.started":"2025-06-24T01:10:22.864143Z","shell.execute_reply":"2025-06-24T01:10:22.872095Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This block evaluates the **CatBoost Ranker** model on the validation set using multiple ranking-specific metrics. Here's a concise explanation without code:\n\n---\n\n### 📈 **1. Predict and Package Results**\n\n* The model makes predictions on the validation set.\n* A DataFrame is built containing:\n\n  * Group IDs (`ranker_id`)\n  * Prediction scores\n  * Ground-truth labels (`selected` = 1 for correct answer)\n\n---\n\n### 🎯 **2. Top-1 Accuracy**\n\n* For each group, the item with the **highest predicted score** is selected.\n* The model’s **Top-1 Accuracy** is the fraction of those predictions where the correct item (`selected == 1`) is at the top.\n\n---\n\n### 🔄 **3. Sigmoid + Log Loss**\n\n* Raw prediction scores are passed through a **sigmoid** function to approximate probabilities.\n* **Log loss** is then calculated on just the top predictions per group—this measures how well-calibrated the model’s confidence is.\n\n---\n\n### 📊 **4. HitRate\\@3**\n\n* Checks if the correct item is **within the top 3 predictions** for each group.\n* Only includes groups with **more than 10 items**, to ensure meaningful ranking diversity.\n* A high HitRate\\@3 indicates good practical performance (e.g., for flight search results).\n\n---\n\n### 📏 **5. Group Size Summary**\n\n* Reports:\n\n  * Average group size in validation\n  * How many groups had more than 10 options\n  * What percentage that represents\n\n---\n\n### ✅ **Why These Metrics Matter**\n\n* **Top-1 Accuracy**: Shows model precision when users only see one recommendation.\n* **HitRate\\@3**: Reflects usefulness in limited-display contexts.\n* **Log Loss**: Captures score calibration, important for confidence-based filtering or ensembling.\n* **Group filtering**: Ensures evaluations reflect real-world conditions with sufficient choice diversity.\n","metadata":{}},{"cell_type":"code","source":"# Evaluate on validation\nval_preds = model.predict(X_val)\nval_df = pd.DataFrame({\n    'ranker_id': groups_val,\n    'pred': val_preds,\n    'selected': y_val\n})\n\n# Get top prediction per group\ntop_preds = val_df.loc[val_df.groupby('ranker_id')['pred'].idxmax()]\ntop_preds['prob'] = sigmoid(top_preds['pred'])\nval_logloss = log_loss(top_preds['selected'], top_preds['prob'])\n\nhitrate_at_3 = calculate_hitrate_at_k(val_df, k=3)\n\n# Additional metrics\nval_accuracy = (top_preds['selected'] == 1).mean()\ngroup_sizes = val_df.groupby('ranker_id').size()\navg_group_size = group_sizes.mean()\n\nprint(f\"HitRate@3 (groups >10):  {hitrate_at_3:.4f}\")\nprint(f\"\\nLogLoss:                 {val_logloss:.4f}\")\nprint(f\"Top-1 Accuracy:          {val_accuracy:.4f}\")\nprint(f\"Groups with >10 options: {(group_sizes > 10).sum()} / {len(group_sizes)} ({(group_sizes > 10).mean():.1%})\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:46:39.658673Z","iopub.execute_input":"2025-06-23T13:46:39.658972Z","iopub.status.idle":"2025-06-23T13:46:41.720345Z","shell.execute_reply.started":"2025-06-23T13:46:39.658945Z","shell.execute_reply":"2025-06-23T13:46:41.715385Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This code finalizes your model’s **test predictions** and prepares a **ranking-style submission file** for evaluation or competition (e.g., Kaggle). Here's a concise explanation:\n\n---\n\n### 📦 **1. Predict on Test Data**\n\n* Uses `CatBoost.Pool` with group IDs and categorical features to structure the test set.\n* The trained model predicts a **ranking score** for each item.\n\n---\n\n### 📊 **2. Prepare Submission Data**\n\n* Builds a new DataFrame with:\n\n  * `Id`: Unique identifier of the test item\n  * `ranker_id`: Group or query ID\n  * `pred_score`: Model’s predicted score\n\n---\n\n### 🏅 **3. Assign Rankings**\n\n* For each group (`ranker_id`), items are **ranked by score (descending)**.\n* The highest score gets `selected = 1`, the next `2`, and so on.\n\n---\n\n### 💾 **4. Save Results**\n\n* Saves the required columns (`Id`, `ranker_id`, `selected`) in **Parquet format** for efficient storage.\n* Optionally, can be saved as CSV if needed.\n\n---\n\n### ✅ **Why It Matters**\n\n* **Ranking integrity is preserved** by grouping before ranking.\n* The submission is ready for platforms that expect groupwise rankings, such as **learning-to-rank competitions**.\n* Efficient file format (Parquet) helps with large-scale evaluation.\n\n","metadata":{}},{"cell_type":"code","source":"# Create test pool and predict\ntest_pool = Pool(X_test, group_id=groups_test, cat_features=cat_features_final)\ntest_preds = model.predict(test_pool)\n\n# Create submission\nsubmission = test[['Id', 'ranker_id']].copy()\nsubmission['pred_score'] = test_preds\n\n# Assign ranks (1 = best option)\nsubmission['selected'] = submission.groupby('ranker_id')['pred_score'].rank(\n    ascending=False, method='first'\n).astype(int)\n\n# Save submission\nsubmission[['Id', 'ranker_id', 'selected']].to_parquet('submission.parquet', index=False)\n#submission[['Id', 'ranker_id', 'selected']].to_csv('submission.csv', index=False)\nprint(f\"Submission saved. Shape: {submission.shape}\")\ndisplay(submission[['Id', 'ranker_id', 'selected']])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-23T13:49:49.665083Z","iopub.execute_input":"2025-06-23T13:49:49.665466Z","iopub.status.idle":"2025-06-23T13:50:31.225319Z","shell.execute_reply.started":"2025-06-23T13:49:49.665432Z","shell.execute_reply":"2025-06-23T13:50:31.218123Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Overall Pipeline\n\nHere’s a high-level overview of the **entire learning-to-rank pipeline** you’ve been working on, focusing on flight ranking as an example:\n\n---\n\n### 1. **Data Preparation & Sampling**\n\n* Start with raw flight offer data including prices, durations, carriers, baggage, etc.\n* If needed, **sample groups** (like queries or users) to reduce dataset size while preserving group integrity.\n* Maintain **group identifiers (`ranker_id`)** to keep items belonging to the same ranking task together.\n\n---\n\n### 2. **Feature Engineering**\n\n* Convert raw time strings into numeric durations.\n* Create **price-related features** (log price, tax ratios, ranks within group).\n* Calculate **duration and segment counts** and encode trip types (one-way, return).\n* Extract **categorical features** such as frequent flyer programs, major carriers, and baggage flags.\n* Add **time-of-day and day-of-week features** for departures and arrivals.\n* Generate **ranking-aware features**, like price rank, cheapest/more expensive flags, and differences from medians.\n* Handle missing values carefully to ensure model robustness.\n\n---\n\n### 3. **Train-Validation Splitting**\n\n* Use a **group-aware splitter** to prevent leakage by ensuring no group appears in both train and validation.\n* This helps simulate real-world generalization to unseen queries or sessions.\n\n---\n\n### 4. **Model Setup**\n\n* Convert datasets into **CatBoost Pools**, including features, labels, group IDs, and categorical feature indicators.\n* Initialize a **CatBoost Ranker** with a ranking-specific loss (e.g., YetiRank), evaluation metrics (e.g., Precision\\@k), and early stopping.\n* Train the model with validation monitoring to avoid overfitting.\n\n---\n\n### 5. **Evaluation**\n\n* Predict on validation set and assess with ranking metrics such as:\n\n  * **HitRate\\@k**: How often relevant items are in the top-k predictions.\n  * **Top-1 Accuracy**: Correct item ranked first.\n  * **Log Loss** on predicted probabilities for calibration insight.\n* Focus evaluation on groups with enough items to ensure meaningful ranking.\n\n---\n\n### 6. **Final Prediction & Submission**\n\n* Predict on the test set using the trained model.\n* Assign ranks per group based on predicted scores.\n* Save results in a submission-ready format preserving group structure.\n\n---\n\n### Summary\n\nThis pipeline emphasizes:\n\n* **Group-aware processing** at every stage.\n* Rich **feature engineering** capturing domain-specific flight info.\n* Usage of **specialized ranking objectives and metrics** to optimize relative ordering rather than absolute predictions.\n* Careful validation and model monitoring to generalize well.\n\n---\n\n\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}