{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.18","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"tpu1vmV38","dataSources":[{"sourceId":105399,"databundleVersionId":12733338,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ✈️Flight Selection Analysis and Prediction — *FlightRank 2025*\n\nThis notebook tackles a **ranking-based recommendation system** challenge applied to the context of **corporate travel**. The goal is to predict which flight a user (business traveler) is most likely to select among multiple available options in a search session.\n\n---\n\n## 🎯Problem and Objective\n\nThis is a **supervised learning problem with a ranking focus**, where observations are grouped by search sessions (`ranker_id`).\n\n- Each `ranker_id` represents a real flight search.\n- Each group contains multiple flight options.\n- Exactly **one** of these options was selected (`selected = 1`).\n\nOur goal is to **train a model that ranks the options correctly** so that the flight chosen by the user appears among the top-ranked options.\n\n---\n\n## 📏Evaluation Metric: HitRate@3\n\nThe official competition metric is **HitRate@3**, which checks whether the flight actually chosen appears among the **top 3 ranked options** for each search group.\n\nThe formula is:\n\n`HitRate@3 = (1 / |Q|) * ∑ 𝟙(rank_i ≤ 3)`\n\nWhere:\n\n- `|Q|` is the number of evaluated search sessions (with more than 10 flights).\n- `rank_i` is the rank assigned by the model to the correct flight in session `i`.\n- `𝟙(rank_i ≤ 3)` equals 1 if the flight is in the top 3, and 0 otherwise.\n\n> **Important:** only sessions with **more than 10 flight options** are considered in the final metric.","metadata":{"_uuid":"ea9825c9-05cf-4c55-a49b-f3da212789db","_cell_guid":"d068ccbd-abb4-4f0f-8acd-e1ee50f2c4e9","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"!pip install lightgbm","metadata":{"_uuid":"f0720e47-ecef-4d29-8636-4b11be46de8a","_cell_guid":"663dbcae-9d63-40bc-acac-36a540711017","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-07-24T06:50:06.014541Z","iopub.execute_input":"2025-07-24T06:50:06.014877Z","iopub.status.idle":"2025-07-24T06:50:11.511102Z","shell.execute_reply.started":"2025-07-24T06:50:06.014854Z","shell.execute_reply":"2025-07-24T06:50:11.506707Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 0. Import Dependencies","metadata":{"_uuid":"e8c8f5bd-b2a8-444f-9673-54d04341ca46","_cell_guid":"492a232e-ff99-4780-9138-ac56743ad294","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"import pandas as pd\nimport os\nimport gc\nimport subprocess\nimport zipfile\nimport matplotlib.pyplot as plt\n\nimport lightgbm as lgb\nimport numpy as np\n\nfrom itertools import product\nfrom sklearn.model_selection import GroupShuffleSplit\nfrom itertools import chain\nfrom sklearn.model_selection import GroupKFold","metadata":{"_uuid":"a1116568-2df1-476c-8302-ff6fa5d4e3db","_cell_guid":"8422a7d7-8169-48d6-b802-83a85beb115d","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:50:18.667982Z","iopub.execute_input":"2025-07-24T06:50:18.668266Z","iopub.status.idle":"2025-07-24T06:50:21.453524Z","shell.execute_reply.started":"2025-07-24T06:50:18.668241Z","shell.execute_reply":"2025-07-24T06:50:21.449135Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# resetting the settings\npd.reset_option('display.max_columns')\n\n# setting the maximum column display option\npd.set_option('display.max_columns', None)","metadata":{"_uuid":"89d9a6eb-6cc9-4b9e-b295-fe3a954ea27d","_cell_guid":"3606ddc8-98fe-4094-a8b0-204c66df47f0","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:50:24.315857Z","iopub.execute_input":"2025-07-24T06:50:24.316297Z","iopub.status.idle":"2025-07-24T06:50:24.326658Z","shell.execute_reply.started":"2025-07-24T06:50:24.316268Z","shell.execute_reply":"2025-07-24T06:50:24.321390Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ⚙️2 .Extraction of Data","metadata":{"_uuid":"79fe0851-55e9-40ea-8b53-a3dfa0e2af59","_cell_guid":"6311d800-44bb-4c1a-b565-bcef7758f5e4","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"\n\ndef prepare_kaggle_dataset():\n    # Set the competition input path (automatically provided by Kaggle)\n    extract_path = \"/kaggle/input/aeroclub-recsys-2025\"\n    \n    # List files in the competition folder\n    print(\"📂 Files available in the dataset folder:\")\n    for root, dirs, files in os.walk(extract_path):\n        for file in files:\n            print(f\"- {file}\")\n    \n    return extract_path  # So you can load files from here\ngc.collect()","metadata":{"_uuid":"3e301ae6-e637-4fa0-b102-ca75865b064d","_cell_guid":"63f4d67d-0d47-4a97-b4ba-2069c31e5c71","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:50:27.325373Z","iopub.execute_input":"2025-07-24T06:50:27.325648Z","iopub.status.idle":"2025-07-24T06:50:27.410532Z","shell.execute_reply.started":"2025-07-24T06:50:27.325624Z","shell.execute_reply":"2025-07-24T06:50:27.407381Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_path = prepare_kaggle_dataset()\n# download_files()\ngc.collect()","metadata":{"_uuid":"55dde75c-9b50-4112-a5d3-d172c2622097","_cell_guid":"a318c544-2681-44d0-8288-f0522098e1fa","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:50:30.201538Z","iopub.execute_input":"2025-07-24T06:50:30.201845Z","iopub.status.idle":"2025-07-24T06:50:30.297977Z","shell.execute_reply.started":"2025-07-24T06:50:30.201818Z","shell.execute_reply":"2025-07-24T06:50:30.293752Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📚 3. Data Loading\n\nIn this step, we load the **training dataset**, available in the `train.parquet` file.\n\nThis dataset contains complete information about **flight search sessions**, including:\n\n- Flight and user identifiers  \n- Company-related data  \n- Route and schedule information  \n- Total price and taxes  \n- Cancellation and rebooking rules  \n- Indication of which flight was selected (`selected = 1`)\n\n> This will be the **main dataset** used for building, training, and validating the recommendation model.","metadata":{"_uuid":"884c1ae6-f504-4d4f-bb93-36aaaffbafb1","_cell_guid":"23f987df-cfc8-43b0-9702-dc12e1fe97d2","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"gc.collect()\n# train = pd.read_parquet(\"data/aeroclub/train.parquet\")\ndata_path = \"/kaggle/input/aeroclub-recsys-2025\"\ntrain = pd.read_parquet(f\"{data_path}/train.parquet\")\ngc.collect()","metadata":{"_uuid":"c028ef24-a188-49cd-9410-838672ab21a3","_cell_guid":"bd11e5ac-4df9-416d-a006-3b0c3f4b3bc7","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:50:34.453530Z","iopub.execute_input":"2025-07-24T06:50:34.453836Z","iopub.status.idle":"2025-07-24T06:50:58.334649Z","shell.execute_reply.started":"2025-07-24T06:50:34.453808Z","shell.execute_reply":"2025-07-24T06:50:58.330588Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def reduce_memory_usage(df):\n    for col in df.columns:\n        col_type = df[col].dtypes\n        \n        if col_type == 'float64':\n            df[col] = pd.to_numeric(df[col], downcast='float')\n        elif col_type == 'int64':\n            df[col] = pd.to_numeric(df[col], downcast='integer')\n        elif col_type == 'object':\n            num_unique = df[col].nunique()\n            num_total = len(df[col])\n            if num_unique / num_total < 0.5:\n                df[col] = df[col].astype('category')\n    \n    return df\ntrain = reduce_memory_usage(train)","metadata":{"_uuid":"6752db61-f609-4631-bd73-764d44b37e1f","_cell_guid":"da6ec350-ebf7-465d-b561-77ef643bb8aa","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:51:04.029965Z","iopub.execute_input":"2025-07-24T06:51:04.030279Z","iopub.status.idle":"2025-07-24T06:52:37.206031Z","shell.execute_reply.started":"2025-07-24T06:51:04.030252Z","shell.execute_reply":"2025-07-24T06:52:37.199814Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"_uuid":"5dcbe36f-381c-4bf9-8311-abe430016e10","_cell_guid":"6eaddd47-bfff-40af-8bd7-4e0bec6021ce","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:52:45.013296Z","iopub.execute_input":"2025-07-24T06:52:45.013584Z","iopub.status.idle":"2025-07-24T06:52:45.109682Z","shell.execute_reply.started":"2025-07-24T06:52:45.013556Z","shell.execute_reply":"2025-07-24T06:52:45.105401Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train_raw = train.copy()\ngc.collect()","metadata":{"_uuid":"79375cd2-0a8e-4e00-bf43-86ddd98ac4e0","_cell_guid":"54140896-fe27-44c3-bbec-25cfc6f49e99","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:53:01.127575Z","iopub.execute_input":"2025-07-24T06:53:01.127906Z","iopub.status.idle":"2025-07-24T06:53:05.750943Z","shell.execute_reply.started":"2025-07-24T06:53:01.127877Z","shell.execute_reply":"2025-07-24T06:53:05.746481Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧹 4. Selection of Relevant Columns\n\nWith the dataset loaded, the next step is to **select only the most relevant columns** for the baseline model.\n\nThe focus is on keeping variables that provide **useful information for flight recommendation**, including:\n\n- Identifiers (`Id`, `ranker_id`, `profileId`, etc.)\n- Passenger and company information\n- Route details, timings, and connections\n- Price data, taxes, and refund/exchange policies\n- Flight segment details (airline, seat, baggage)\n- Target variable: `selected` (indicates the flight chosen by the user)\n\n> This step reduces dimensionality and improves model performance by eliminating irrelevant or redundant columns.\n\n---\n\n### Optional Sampling for Prototyping\n\nDuring early experiments, it is common to work with a **reduced sample of the data** to speed up iteration. For that, we created a helper function that allows:\n\n- Selecting only the desired columns  \n- Optionally limiting the number of rows loaded\n\n### Examples:\n\n```python\n# Load only a sample (1 million rows)\ndf_train = load_subset(df_train_raw, columns_to_keep, max_rows=1_000_000)\n\n# Load the full dataset (all rows)\ndf_train = load_subset(df_train_raw, columns_to_keep)","metadata":{"_uuid":"a926d862-c0ba-4124-9992-361d4186099a","_cell_guid":"4e24e9fe-68bd-4134-aa33-cd8359fe0c11","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"\n# Define the columns you want to keep\n\ncolumns_to_keep = [\n    # Identifiers\n    'Id',  # num\n    'ranker_id', \n    'profileId', \n    'companyID',\n    \n    # User info\n    'sex', 'nationality', 'frequentFlyer', 'isVip', 'bySelf', 'isAccess3D',\n\n    # Company info\n    'corporateTariffCode',\n\n    # Search & route\n    'searchRoute', 'requestDate',\n\n    # Pricing\n    'totalPrice', 'taxes',\n\n    # Flight timing\n    'legs0_departureAt', 'legs0_arrivalAt', 'legs0_duration',\n    'legs1_departureAt', 'legs1_arrivalAt', 'legs1_duration',\n\n    # Segment-level info (só do segmento 0 da ida para simplificar no baseline)\n    'legs0_segments0_departureFrom_airport_iata',\n    'legs0_segments0_arrivalTo_airport_iata',\n    'legs0_segments0_arrivalTo_airport_city_iata',\n    'legs0_segments0_marketingCarrier_code',\n    'legs0_segments0_operatingCarrier_code',\n    'legs0_segments0_aircraft_code',\n    'legs0_segments0_flightNumber',\n    'legs0_segments0_duration',\n    'legs0_segments0_baggageAllowance_quantity',\n    'legs0_segments0_baggageAllowance_weightMeasurementType',\n    'legs0_segments0_cabinClass',\n    'legs0_segments0_seatsAvailable', \n    'legs0_segments1_departureFrom_airport_iata',\n    'legs0_segments2_departureFrom_airport_iata',\n    'legs0_segments3_departureFrom_airport_iata',\n\n    # Cancellation & exchange rules\n    'miniRules0_monetaryAmount', 'miniRules0_percentage', 'miniRules0_statusInfos',\n    'miniRules1_monetaryAmount', 'miniRules1_percentage', 'miniRules1_statusInfos',\n\n    # Pricing policy\n    'pricingInfo_isAccessTP', 'pricingInfo_passengerCount',\n\n    # Target\n    'selected'\n]\n\n# Filter the data for the baseline\ndef load_subset(df, columns,  max_rows=None):\n    if max_rows:\n        return df[columns].iloc[:max_rows].copy()\n    else:\n        return df[columns].copy()\n\n# Example of usage\ndf_train = load_subset(df_train_raw, columns_to_keep, max_rows=1_000_000) # ONLY 1M\n\n#############################          IMPORTANT      ########################################\n#############################          IMPORTANT      ########################################\n#############################          IMPORTANT      ########################################\n#############################          IMPORTANT      ########################################\n#df_train = load_subset(df_train_raw, columns_to_keep) # ALL REGISTERS","metadata":{"_uuid":"220a93d5-2977-473c-9167-5b9ca3141d5b","_cell_guid":"7f90c7bf-92f9-480d-a3c0-d437e461fd8c","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:53:09.287313Z","iopub.execute_input":"2025-07-24T06:53:09.287616Z","iopub.status.idle":"2025-07-24T06:53:10.342472Z","shell.execute_reply.started":"2025-07-24T06:53:09.287588Z","shell.execute_reply":"2025-07-24T06:53:10.338138Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🛠️ 5. Feature Engineering\n\nIn this step, we transform raw columns into more informative, consistent, and suitable variables for use in machine learning models.\n\nThe features will be built in subtopics, organized by type of transformation.\n\n---\n\n### 🧾 5.1 Data Type Correction (`dtypes`)\n\nThe first step is to ensure that data types are correct and optimized.\n\n- Categorical columns initially encoded as `category` are analyzed and converted to:\n  - `int` or `float`, when possible  \n  - `bool`, if they contain only logical values  \n  - `category`, in all other cases\n\n- The `nationality` column, which arrives as an integer, is converted to `string` to preserve its categorical meaning.\n\n> This standardization is essential to avoid errors and ensure that the model interprets variables correctly.","metadata":{"_uuid":"8e64dda9-df9b-4670-84e1-e85eeb5ae43d","_cell_guid":"4cad5c11-399d-40e0-acc2-2723160e5784","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def fix_column_types(df):\n    df_fixed = df.copy()\n    for col in df.columns:\n        if isinstance(df[col].dtype, pd.CategoricalDtype):\n            # Try converting to numeric type\n            try:\n                df_fixed[col] = pd.to_numeric(df[col])\n            except:\n                # If not numeric, try boolean\n                unique_vals = df[col].dropna().unique()\n                if set(unique_vals) <= {True, False}:\n                    df_fixed[col] = df[col].astype(bool)\n                else:\n                    df_fixed[col] = df[col].astype(str)\n    return df_fixed\n\ndf_train = fix_column_types(df_train)\n\n# Adjust nationality (currently in int format)\ndf_train[\"nationality\"] = df_train[\"nationality\"].astype(\"str\")\n\n# Convert companyID to category type\ndf_train['companyID'] = df_train['companyID'].astype('category')\n\n# Check result\ndf_train.dtypes","metadata":{"_uuid":"5d3bc738-602b-407a-9f06-713944902299","_cell_guid":"0ce63804-a21b-46bc-b0fc-56b152caf22e","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:53:22.520493Z","iopub.execute_input":"2025-07-24T06:53:22.520774Z","iopub.status.idle":"2025-07-24T06:53:26.560867Z","shell.execute_reply.started":"2025-07-24T06:53:22.520749Z","shell.execute_reply":"2025-07-24T06:53:26.554768Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 5.2 Feature Engineering","metadata":{"_uuid":"97e33663-defe-42fe-9912-40150bd98afb","_cell_guid":"e3336c70-a513-41eb-84ac-7a9f50a78e22","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def feature_engineer(df):\n    \"\"\"\n    Engineers a comprehensive set of features for the flight ranking model.\n    \"\"\"\n    df = df.copy()\n\n    # 1. Time-Based Features\n    for col in ['requestDate', 'legs0_departureAt', 'legs0_arrivalAt', 'legs1_departureAt', 'legs1_arrivalAt']:\n        df[col] = pd.to_datetime(df[col], errors='coerce')\n\n    df['request_dayofweek'] = df['requestDate'].dt.dayofweek\n    df['request_hour'] = df['requestDate'].dt.hour\n    df['departure_hour'] = df['legs0_departureAt'].dt.hour\n    df['is_weekend'] = df['legs0_departureAt'].dt.dayofweek.isin([5, 6]).astype(int)\n    df['is_business_hours'] = ((df['departure_hour'] >= 8) & (df['departure_hour'] <= 18)).astype(int)\n    df['days_to_departure'] = (df['legs0_departureAt'] - df['requestDate']).dt.days\n\n    # 2. Duration Conversion (String to Hours)\n    def to_hours(duration_str):\n        if pd.isna(duration_str):\n            return np.nan\n        try:\n            parts = str(duration_str).split(':')\n            hours = int(parts[0]) + int(parts[1]) / 60\n            return hours\n        except:\n            return np.nan\n\n    df['legs0_duration_hr'] = df['legs0_duration'].apply(to_hours)\n    df['legs1_duration_hr'] = df['legs1_duration'].apply(to_hours)\n    df['total_duration'] = df['legs0_duration_hr'].fillna(0) + df['legs1_duration_hr'].fillna(0)\n\n    # 3. Segment & Connection Features\n    stop_cols = [c for c in df.columns if 'segments' in c and 'departureFrom' in c]\n    df['num_stops'] = df[stop_cols].notna().sum(axis=1) - 1  # Subtract 1 for the origin\n    df['is_direct_flight'] = (df['num_stops'] == 0).astype(int)\n\n    # 4. Interaction and Ratio Features\n    df['tax_ratio'] = df['taxes'] / (df['totalPrice'] + 1e-6)\n    df['price_per_duration'] = df['totalPrice'] / (df['total_duration'] + 1e-6)\n    df['price_to_stops_ratio'] = df['totalPrice'] / (df['num_stops'] + 1)\n\n    # 5. Group-wise (Ranker ID) Features\n    group_features = ['totalPrice', 'total_duration', 'num_stops']\n    for feature in group_features:\n        df[f'{feature}_avg_ranker'] = df.groupby('ranker_id')[feature].transform('mean')\n        df[f'{feature}_max_ranker'] = df.groupby('ranker_id')[feature].transform('max')\n        df[f'{feature}_min_ranker'] = df.groupby('ranker_id')[feature].transform('min')\n        df[f'{feature}_std_ranker'] = df.groupby('ranker_id')[feature].transform('std')\n\n        df[f'{feature}_diff_from_avg'] = df[feature] - df[f'{feature}_avg_ranker']\n        df[f'{feature}_ratio_to_avg'] = df[feature] / (df[f'{feature}_avg_ranker'] + 1e-6)\n\n    # 6. Categorical and Boolean Handling\n    bool_cols = ['isVip', 'bySelf', 'isAccess3D', 'pricingInfo_isAccessTP']\n    for col in bool_cols:\n        df[col] = df[col].astype(bool)\n\n    cat_cols = [\n        'companyID', 'sex', 'nationality', 'corporateTariffCode',\n        'legs0_segments0_departureFrom_airport_iata', 'legs0_segments0_arrivalTo_airport_iata',\n        'legs0_segments0_marketingCarrier_code', 'legs0_segments0_operatingCarrier_code',\n        'legs0_segments0_cabinClass'\n    ]\n    for col in cat_cols:\n        df[col] = pd.factorize(df[col])[0]\n\n    # 7. Clean up and Final Touches\n    df = df.drop(columns=[\n        'requestDate', 'legs0_departureAt', 'legs0_arrivalAt', 'legs1_departureAt', 'legs1_arrivalAt',\n        'legs0_duration', 'legs1_duration', 'frequentFlyer', 'searchRoute',\n        'legs0_segments0_duration', 'legs0_segments1_departureFrom_airport_iata',\n        'legs0_segments2_departureFrom_airport_iata', 'legs0_segments3_departureFrom_airport_iata'\n    ], errors='ignore')\n\n    # Fill NaNs created during feature engineering\n    df.fillna(-1, inplace=True)\n\n    return df","metadata":{"_uuid":"598dcbcb-6adf-4f02-9617-23cb0c4a56bb","_cell_guid":"219810be-93c0-40e2-bdd1-18de2dbc0cc8","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:53:41.381435Z","iopub.execute_input":"2025-07-24T06:53:41.381745Z","iopub.status.idle":"2025-07-24T06:53:41.401699Z","shell.execute_reply.started":"2025-07-24T06:53:41.381717Z","shell.execute_reply":"2025-07-24T06:53:41.396392Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Apply the feature engineering on training data sample(1,000,000)","metadata":{"_uuid":"1b19b795-ecef-42d7-88b8-6026f3a5230d","_cell_guid":"1de70e7f-c347-42f4-a2db-063cf353f040","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"\nprint(\"Starting feature engineering for the training set...\")\ndf_train_fe = feature_engineer(df_train)\nprint(\"Feature engineering complete.\")\ngc.collect()","metadata":{"_uuid":"14993f89-aea7-4275-bfb1-4d66f9c4b71d","_cell_guid":"d71689bb-cb90-4403-b699-879f656094e6","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:53:47.556829Z","iopub.execute_input":"2025-07-24T06:53:47.557130Z","iopub.status.idle":"2025-07-24T06:53:54.561368Z","shell.execute_reply.started":"2025-07-24T06:53:47.557105Z","shell.execute_reply":"2025-07-24T06:53:54.557209Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Testing Data Loading and memory optimization","metadata":{"_uuid":"335a1a7c-8944-44bc-9da8-c5e0ed80ffca","_cell_guid":"f878ecd5-a97a-4379-8bc0-73d303ae038c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"test = pd.read_parquet(f\"{data_path}/test.parquet\")\ntest = reduce_memory_usage(test)\n\n\ndf_test_raw = test.copy()\nprint(f\"Full test data loaded with {len(df_test_raw)} rows.\")\nprint(\"\\nStarting feature engineering for the full test set...\")\n\ngc.collect()\n","metadata":{"_uuid":"314b676c-7478-44a9-a1c4-e05b1c8012a6","_cell_guid":"b474fba0-595f-48fd-95b4-be5b5e731c6e","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T06:58:34.586429Z","iopub.execute_input":"2025-07-24T06:58:34.586802Z","iopub.status.idle":"2025-07-24T06:59:22.758979Z","shell.execute_reply.started":"2025-07-24T06:58:34.586770Z","shell.execute_reply":"2025-07-24T06:59:22.753517Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Fix test column data types \n## Apply the feature engineering on test data","metadata":{"_uuid":"646954ee-f0ce-4afa-8b12-a33f3c546536","_cell_guid":"ba352066-ce91-4903-97d7-ae0336ed09eb","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"df_test = fix_column_types(df_test_raw)\n\n# Adjust nationality (currently in int format)\ndf_test[\"nationality\"] = df_test[\"nationality\"].astype(\"str\")\n\n# Convert companyID to category type\ndf_test['companyID'] = df_test['companyID'].astype('category')\n\n# Check result\ndf_test.dtypes\n\ndf_test_fe = feature_engineer(df_test)\nprint(\"Feature engineering on the full test set complete.\")\nprint(f\"Processed test data has {len(df_test_fe)} rows.\")\ngc.collect()","metadata":{"_uuid":"62f792e6-e2fe-4b6a-9a94-1e05429e49b6","_cell_guid":"fd2bdae7-f80b-4d83-9a5f-90ab97c17d98","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:02:43.317583Z","iopub.execute_input":"2025-07-24T07:02:43.317988Z","iopub.status.idle":"2025-07-24T07:05:02.809764Z","shell.execute_reply.started":"2025-07-24T07:02:43.317958Z","shell.execute_reply":"2025-07-24T07:05:02.805594Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧪 6. Training Preparation\n\nWith all features processed, the next step is to prepare the data for training the ranking model with LightGBM.\n\n---\n\n### 🎯 6.1 Target and Group Definition\n\n- **`target_col`**: target variable indicating whether the flight was selected (`selected = 1`).\n- **`group_col`**: identifies each flight search session (`ranker_id`), used to properly group options in the ranking model.\n\n---\n\n### 🧮 6.2 Feature Organization\n\nFeatures are divided into three types:\n\n- 🔢 **Numerical (`numeric_cols`)**: continuous values like price, duration, baggage count, etc.\n- 🏷️ **Categorical (`categorical_cols`)**: variables representing codes, airports, airlines, etc.\n- ✅ **Boolean (`boolean_cols`)**: indicator variables (e.g., `isVip`, `ida_fds`, `hasFrequentFlyer`, etc.)\n\nThese lists are combined into the final `features` variable, which will be used as input for the model.\n\n---\n\n### 🧪 6.3 Train/Validation Split\n\n`GroupShuffleSplit` is used to perform the **split while respecting groups (`ranker_id`)**, ensuring that all options from the same search session appear **either in training or in validation**, but not both.\n\n---\n\n### 📦 6.4 Dataset Construction for LightGBM\n\nThe `train_dataset` and `val_dataset` objects are created, which are optimized LightGBM structures for ranking:\n\n- Include the data (`X_train`, `X_val`) and targets (`y_train`, `y_val`)\n- Receive the list of categorical columns\n- Incorporate the groups (`group=...`) required for **supervised ranking**\n- Define the `max_bin` parameter, which controls discretization of continuous variables (used to speed up training and allow GPU usage)\n\n> This structure is essential for using the **`lambdarank` objective**, as the model needs to understand the comparison groups.","metadata":{"_uuid":"0733fdf7-82a1-49af-b3a1-8552fff132e5","_cell_guid":"7aa4a8a4-09f3-462f-8885-3f3aa68f195b","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"# --- Target and group column","metadata":{"_uuid":"1d6cabb8-f076-4a18-a19a-fb63dc2be413","_cell_guid":"c31aafe3-c747-4cd1-99a9-4495b3e37bcc","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# --- Target and group column\ntarget_col = \"selected\"\ngroup_col = \"ranker_id\"","metadata":{"_uuid":"4682e1c4-fc9c-48b9-b2c6-87a83a318d7e","_cell_guid":"a7f48df8-d32a-4b50-8817-743589e0b3f3","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:05:44.299250Z","iopub.execute_input":"2025-07-24T07:05:44.299550Z","iopub.status.idle":"2025-07-24T07:05:44.309414Z","shell.execute_reply.started":"2025-07-24T07:05:44.299524Z","shell.execute_reply":"2025-07-24T07:05:44.304485Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## --- Define features for the model","metadata":{"_uuid":"82a0d131-04cb-42d3-9c32-8d2c53561a4e","_cell_guid":"f361729e-8cce-42f4-bf7e-241d7e0c8752","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"features = [col for col in df_train_fe.columns if col not in ['Id', 'ranker_id', 'profileId', 'selected']]\ncategorical_features = [\n'companyID', 'sex', 'nationality', 'corporateTariffCode',\n'legs0_segments0_departureFrom_airport_iata', 'legs0_segments0_arrivalTo_airport_iata',\n'legs0_segments0_marketingCarrier_code', 'legs0_segments0_operatingCarrier_code',\n'legs0_segments0_cabinClass','legs0_segments0_arrivalTo_airport_city_iata',  \n'legs0_segments0_aircraft_code' \n]","metadata":{"_uuid":"1053d985-09d7-4b57-a6e1-76109f03cccd","_cell_guid":"6e476dd9-463d-4c63-b207-70acbfb5990d","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:05:46.365825Z","iopub.execute_input":"2025-07-24T07:05:46.366202Z","iopub.status.idle":"2025-07-24T07:05:46.377105Z","shell.execute_reply.started":"2025-07-24T07:05:46.366174Z","shell.execute_reply":"2025-07-24T07:05:46.371619Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Ensure all categorical features are actually in the list of features","metadata":{"_uuid":"cae5087b-8bb4-4a54-bfa7-1a4b9476a002","_cell_guid":"32244d77-3af0-4ae4-96da-938eb4096011","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"categorical_features = [f for f in categorical_features if f in features]\n# for col in categorical_features:\n#     df_train_fe[col] = pd_train_fe.factorize(df[col])[0]","metadata":{"_uuid":"356b805f-22d4-4464-aed4-ecddb7b62cec","_cell_guid":"2f912814-87e1-41f1-ad1f-c18a420a9638","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:05:48.642129Z","iopub.execute_input":"2025-07-24T07:05:48.642430Z","iopub.status.idle":"2025-07-24T07:05:48.653249Z","shell.execute_reply.started":"2025-07-24T07:05:48.642406Z","shell.execute_reply":"2025-07-24T07:05:48.647748Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n## --- Convert categorical columns to 'category' dtype for LightGBM","metadata":{"_uuid":"a29204b5-030a-4018-bfdb-1121feb82753","_cell_guid":"e372f866-c707-4ca7-b4aa-53226341d624","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"for col in categorical_features:\n    df_train_fe[col] = df_train_fe[col].astype(\"category\")\n    df_test_fe[col] = df_test_fe[col].astype(\"category\")","metadata":{"_uuid":"edb6668e-7443-4b44-a775-2e003e431506","_cell_guid":"243926bf-703e-4261-a350-cf5b7fed2f26","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:05:50.791493Z","iopub.execute_input":"2025-07-24T07:05:50.791783Z","iopub.status.idle":"2025-07-24T07:05:52.371710Z","shell.execute_reply.started":"2025-07-24T07:05:50.791760Z","shell.execute_reply":"2025-07-24T07:05:52.366558Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## --- Split by group (ranker_id)","metadata":{"_uuid":"f56365e4-64a2-43b1-8b80-9dbecc503fa6","_cell_guid":"4d0cdc59-cd79-4921-aa9c-25ba16d5d77e","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"gss = GroupShuffleSplit(n_splits=1, test_size=0.2, random_state=42)\ntrain_idx, val_idx = next(gss.split(df_train_fe, groups=df_train_fe[\"ranker_id\"]))\ndf_train_split = df_train_fe.iloc[train_idx].copy()\ndf_val = df_train_fe.iloc[val_idx].copy()","metadata":{"_uuid":"ec068e99-2970-43b8-be9f-0ed019817be2","_cell_guid":"33b28313-09e8-41be-9d85-0a27453a5758","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:05:57.285247Z","iopub.execute_input":"2025-07-24T07:05:57.285568Z","iopub.status.idle":"2025-07-24T07:05:58.597173Z","shell.execute_reply.started":"2025-07-24T07:05:57.285540Z","shell.execute_reply":"2025-07-24T07:05:58.592339Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## --- Features and targets","metadata":{"_uuid":"5f555549-86cf-4997-955c-93986cdb6135","_cell_guid":"d6f8901a-d4ce-4172-b500-5e0276ab79f0","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"X_train = df_train_split[features]\ny_train = df_train_split[target_col]\ngroups_train = df_train_split.groupby(group_col).size().to_numpy()\nX_val = df_val[features]\ny_val = df_val[target_col]\ngroups_val = df_val.groupby(group_col).size().to_numpy()\ndataset_params = {\n\"max_bin\": 63\n}","metadata":{"_uuid":"021edbfe-aa72-4328-90df-ece9c1d8ea5b","_cell_guid":"6a6d1802-a1cd-4ff9-bd0d-0a09da7ac2f6","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:06:00.845852Z","iopub.execute_input":"2025-07-24T07:06:00.846195Z","iopub.status.idle":"2025-07-24T07:06:01.039491Z","shell.execute_reply.started":"2025-07-24T07:06:00.846165Z","shell.execute_reply":"2025-07-24T07:06:01.034723Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## --- Dataset creation","metadata":{"_uuid":"594444e3-2397-47b7-913a-0248dc45f5e9","_cell_guid":"76c159c5-f298-49c6-893a-a37c737dd373","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"train_dataset = lgb.Dataset(\nX_train,\nlabel=y_train,\ngroup=groups_train,\ncategorical_feature=categorical_features,\nparams=dataset_params\n)","metadata":{"_uuid":"8f7c80ff-4ab5-4473-b47d-1f8429b8a351","_cell_guid":"894e866c-b69f-42eb-aef9-05490c990a9c","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:06:03.110641Z","iopub.execute_input":"2025-07-24T07:06:03.110970Z","iopub.status.idle":"2025-07-24T07:06:03.121872Z","shell.execute_reply.started":"2025-07-24T07:06:03.110943Z","shell.execute_reply":"2025-07-24T07:06:03.116266Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"val_dataset = lgb.Dataset(\nX_val,\nlabel=y_val,\ngroup=groups_val,\ncategorical_feature=categorical_features,\nreference=train_dataset,\nparams=dataset_params\n)\nprint(\"Train and validation datasets created successfully.\")","metadata":{"_uuid":"cae4b036-4656-45ff-a01b-65e25c1f0d25","_cell_guid":"98f379ab-52b0-45d3-82e3-7b210d9e4de9","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:06:32.417903Z","iopub.execute_input":"2025-07-24T07:06:32.418265Z","iopub.status.idle":"2025-07-24T07:06:32.429648Z","shell.execute_reply.started":"2025-07-24T07:06:32.418235Z","shell.execute_reply":"2025-07-24T07:06:32.423744Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧮 7. Hyperparameter Search with LightGBM\n\nBefore training the final model, we perform a **manual Grid Search** to identify the best combination of hyperparameters for the `Lambdarank` model.\n\n---\n\n### 🔧 7.1 Parameters Tested\n\nThe search is conducted over the following hyperparameters:\n\n- `learning_rate`: learning rate (e.g., 0.05)\n- `num_leaves`: tree complexity (e.g., 63, 127)\n- `min_data_in_leaf`: regularization via minimum samples per leaf (e.g., 50, 70, 100)\n\nAll possible combinations of these values are tested using `itertools.product`.\n\n---\n\n### 🧪 7.2 Training and Validation\n\nFor each parameter combination:\n\n1. The LightGBM model is trained with:\n   - Objective: `lambdarank`\n   - Metric: `ndcg@3`\n   - Early stopping after 50 rounds without improvement\n\n2. The model's performance is evaluated based on the **best NDCG@3** score achieved on the validation set.\n\n3. The best model and parameter set are stored.\n\n---\n\n### ✅ 7.3 Search Result\n\nAt the end of the search:\n\n- The **best parameter combination** is displayed\n- The **NDCG@3 score** is reported\n- A **validation prediction** is performed using the best model\n- The **top-1 accuracy** is calculated, i.e., the fraction of sessions where the correct flight was ranked first by the model\n\n> This evaluation serves as a practical check of the recommendation quality before the final training on the full dataset.","metadata":{"_uuid":"303d3022-eca2-4f14-a5b0-c65aaa8196d8","_cell_guid":"a132c1b3-17ec-4380-a9a4-1085ea4454a9","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"param_grid = {\n'learning_rate': [0.05, 0.1],\n'num_leaves': [63, 127],\n\n'min_child_samples':[50, 70, 100]\n\n}","metadata":{"_uuid":"188efbeb-b24e-4d94-a3bf-e78c9f723e45","_cell_guid":"80b779e4-f9d1-408a-87db-d39c4eb50418","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:06:37.583541Z","iopub.execute_input":"2025-07-24T07:06:37.583829Z","iopub.status.idle":"2025-07-24T07:06:37.593990Z","shell.execute_reply.started":"2025-07-24T07:06:37.583805Z","shell.execute_reply":"2025-07-24T07:06:37.589286Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Generate all parameter combinations","metadata":{"_uuid":"31c5a7d6-2086-4eac-8ca6-327afbb6ea0b","_cell_guid":"f30fd591-ffd2-4716-a60b-24c48f940588","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"param_combinations = list(product(*param_grid.values()))\nparam_keys = list(param_grid.keys())\nbest_score = -1\nbest_model = None\nbest_params = None\nfor combo in param_combinations:\n    param_set = dict(zip(param_keys, combo))\nprint(f\"Training with: {param_set}\")","metadata":{"_uuid":"4dbd12d3-0ee0-4f0d-8c37-94fb2eb6b6ae","_cell_guid":"0e72a45f-a64c-4ff7-836f-9c96e4828d53","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:06:39.760221Z","iopub.execute_input":"2025-07-24T07:06:39.760542Z","iopub.status.idle":"2025-07-24T07:06:39.770925Z","shell.execute_reply.started":"2025-07-24T07:06:39.760514Z","shell.execute_reply":"2025-07-24T07:06:39.766318Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"params = {\n    \"objective\": \"lambdarank\",\n    \"metric\": \"ndcg\",\n    \"ndcg_eval_at\":[3],\n    \"boosting_type\": \"gbdt\",\n    \"feature_fraction\": 0.8,\n    \"bagging_fraction\": 0.8,\n    \"bagging_freq\": 1,\n    \"seed\": 42,\n    \"verbosity\": -1,\n    \"num_threads\": -1, # Use all available threads\n    **param_set\n}","metadata":{"_uuid":"5378e915-aa36-4173-b3e1-c36121d1c6cc","_cell_guid":"e3626e86-ffcb-4f98-a9e1-ff25786a0e1f","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:06:41.559502Z","iopub.execute_input":"2025-07-24T07:06:41.559818Z","iopub.status.idle":"2025-07-24T07:06:41.569934Z","shell.execute_reply.started":"2025-07-24T07:06:41.559790Z","shell.execute_reply":"2025-07-24T07:06:41.565632Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = lgb.train(\n    params,\n    train_dataset,\n    valid_sets=[val_dataset],\n    valid_names=[\"valid\"],\n    num_boost_round=1000,\n    callbacks=[lgb.early_stopping(stopping_rounds=50, verbose=False)],\n)\n\nscore = model.best_score[\"valid\"][\"ndcg@3\"]\nprint(f\"  -> Achieved NDCG@3: {score:.5f}\")","metadata":{"_uuid":"1ccc326f-d455-41f9-8170-b862e00638b0","_cell_guid":"f227bacb-6b1f-44d7-85d6-bdb922fcb775","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:06:43.317059Z","iopub.execute_input":"2025-07-24T07:06:43.317368Z","iopub.status.idle":"2025-07-24T07:07:22.236399Z","shell.execute_reply.started":"2025-07-24T07:06:43.317343Z","shell.execute_reply":"2025-07-24T07:07:22.231003Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if score > best_score:\n    best_score = score\n    best_model = model\n    best_params = param_set\n\nprint(\"\\n✅ Best combination found:\")\nprint(f\"   Parameters: {best_params}\")\nprint(f\"   Best NDCG@3 on validation set: {best_score:.5f}\")","metadata":{"_uuid":"70893637-254c-41e2-aa30-42d4e250a638","_cell_guid":"71fffa13-b5f1-4510-b1a0-87b7712ffca6","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:07:26.933517Z","iopub.execute_input":"2025-07-24T07:07:26.933901Z","iopub.status.idle":"2025-07-24T07:07:26.946348Z","shell.execute_reply.started":"2025-07-24T07:07:26.933868Z","shell.execute_reply":"2025-07-24T07:07:26.940583Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Final Training with the Entire Dataset","metadata":{"_uuid":"7a29117c-476d-423c-82dc-158b6a5125aa","_cell_guid":"e9ccbc16-0510-487c-8b21-b2ac00d0f780","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":" After validating the model and identifying the best hyperparameter configuration, we proceed with the final training using 100% of the training data. This ensures the model learns from all available information before making predictions on the unseen test set","metadata":{"_uuid":"b1f8848c-3ccc-49ef-958e-818b640951f3","_cell_guid":"0d672720-a601-4094-81fa-cac12bceaab6","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"### 8.1 Full Dataset Preparation","metadata":{"_uuid":"33e24a7e-d396-4661-8e3a-d4e0e02fc25f","_cell_guid":"f63ae88c-3bc9-46bd-b53f-90f42df6e588","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"We now construct a single LightGBM Dataset object using the entire feature-engineered training set (df_train_fe).\n- X_full: All features from the complete training data.\n- y_full: The corresponding selected target labels.\n- groups_full: The group structure (ranker_id) for all training sessions.","metadata":{"_uuid":"6440dbd9-b29c-4f3f-bb60-2bc55db3eb78","_cell_guid":"47645909-1aab-4053-b6f9-aad0c43bf2fe","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T05:17:39.994438Z","iopub.execute_input":"2025-07-24T05:17:39.994841Z","iopub.status.idle":"2025-07-24T05:17:40.008478Z","shell.execute_reply.started":"2025-07-24T05:17:39.994812Z","shell.execute_reply":"2025-07-24T05:17:40.003820Z"},"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"### 8.2 Using best_iteration for Optimal Training","metadata":{"_uuid":"1c642f97-6133-4d55-8636-0a3654fc1d1e","_cell_guid":"6c15deaa-2cbf-4014-bd0f-4b725c1e7491","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"The validation process with early_stopping provided a crucial piece of information: model.best_iteration. This value represents the optimal number of boosting rounds where the model achieved its peak performance on the validation set before starting to overfit.\n\nBy training the final model on the full dataset for exactly this number of rounds, we leverage the insights from our validation to create a robust and well-generalized model.","metadata":{"_uuid":"de5ef370-b74b-49eb-9567-f28a3bdb01ab","_cell_guid":"4796d145-6c30-4696-bcf4-046f3a97a081","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"### --- Prepare the full dataset for final training","metadata":{"_uuid":"1b7415ef-f290-486f-b556-2ec0ca809e08","_cell_guid":"e40f3731-467e-4e79-b57a-24f691c6f430","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T05:18:52.023730Z","iopub.execute_input":"2025-07-24T05:18:52.024031Z","iopub.status.idle":"2025-07-24T05:18:52.033272Z","shell.execute_reply.started":"2025-07-24T05:18:52.024008Z","shell.execute_reply":"2025-07-24T05:18:52.028824Z"},"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"X_full = df_train_fe[features]\ny_full = df_train_fe[target_col]\ngroups_full = df_train_fe.groupby(group_col).size().to_numpy()\nfull_dataset = lgb.Dataset(\nX_full,\ny_full,\ngroup=groups_full,\ncategorical_feature=categorical_features,\nparams=dataset_params\n)","metadata":{"_uuid":"386e777e-4df0-40b9-a2e2-77e1a41c55da","_cell_guid":"850c6b9a-8c2c-4546-9030-f153014a4c7e","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:07:31.518654Z","iopub.execute_input":"2025-07-24T07:07:31.519013Z","iopub.status.idle":"2025-07-24T07:07:31.834108Z","shell.execute_reply.started":"2025-07-24T07:07:31.518984Z","shell.execute_reply":"2025-07-24T07:07:31.829349Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### --- Final model parameters","metadata":{"_uuid":"4f3a489d-84b7-4ccd-bca8-e3eb3c2a4b8e","_cell_guid":"d29bce5a-2d9a-4753-9acb-7f3aa1b594ae","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"final_params = {\n\"objective\": \"lambdarank\",\n\"metric\": \"ndcg\",\n\"ndcg_eval_at\": [3],\n\"boosting_type\": \"gbdt\",\n\"feature_fraction\": 0.8,\n\"bagging_fraction\": 0.8,\n\"bagging_freq\": 1,\n\"seed\": 42,\n\"verbosity\": -1,\n\"num_threads\": -1,\n**best_params  # Use the best parameters found in the grid search\n}","metadata":{"_uuid":"b268d4a5-cbde-48d4-8170-e57b56097fa1","_cell_guid":"19676f38-fbbc-4643-8ffa-b43460507d22","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:07:33.549862Z","iopub.execute_input":"2025-07-24T07:07:33.550234Z","iopub.status.idle":"2025-07-24T07:07:33.561819Z","shell.execute_reply.started":"2025-07-24T07:07:33.550205Z","shell.execute_reply":"2025-07-24T07:07:33.556329Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## --- Train the final model\nWe use the ideal number of iterations found during validation with early stopping.","metadata":{"_uuid":"86609e4f-bd6a-4e5a-9b84-8efb18dfdef5","_cell_guid":"a46366bb-20cd-49ff-8c64-5fd228911117","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"print(f\"Training final model with best parameters for {best_model.best_iteration} rounds...\")\nfinal_model = lgb.train(\nfinal_params,\nfull_dataset,\nnum_boost_round=best_model.best_iteration\n)\nprint(\"✅ Final model training complete.\")","metadata":{"_uuid":"e31ea36d-b891-4147-af1f-9d1b11b8bdd6","_cell_guid":"4aa5a3fc-b4aa-434d-98c8-b75dee9b59f6","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:07:35.607852Z","iopub.execute_input":"2025-07-24T07:07:35.608246Z","iopub.status.idle":"2025-07-24T07:08:10.874166Z","shell.execute_reply.started":"2025-07-24T07:07:35.608214Z","shell.execute_reply":"2025-07-24T07:08:10.866865Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 9. Submission Generation\nWith the final model trained on the entire dataset, the last step is to generate the submission file for the competition.","metadata":{"_uuid":"d125b292-7db8-4273-a0ff-1a84292eab71","_cell_guid":"7b6a4cbb-962f-482f-a60d-c58745a85694","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"### Process:\n1. Load Test Data: The feature-engineered test set (df_test_fe) is used.\n2. Predict Scores: The trained final_model predicts a relevance score for each flight option in the test set.\n3. Rank Flights: Within each ranker_id group, the flights are sorted in descending order based on their predicted scores.\n4. Assign Ranks: A rank (from 1 upwards) is assigned to each flight within its group. This becomes the selected column for the submission.\n5. Format and Save: The final submission file (submission.csv) is created with the required Id, ranker_id, and selected columns.","metadata":{"_uuid":"eced0106-0453-40ff-98d1-784f2810d47d","_cell_guid":"033eeeb6-c8ab-4e82-a496-21ebe4c23ec4","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"### --- Predict on the test set","metadata":{"_uuid":"ce3162f1-d0b6-40a7-b575-ff3384db6c2c","_cell_guid":"5926e183-0ffa-4957-bc4c-ad0bbcbdb4f3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"\nprint(\"Predicting on the test set...\")\nX_test = df_test_fe[features]\ndf_test_fe['y_pred'] = final_model.predict(X_test)","metadata":{"_uuid":"8d3c851b-1d62-4a87-bb41-335f17dbafd4","_cell_guid":"c8cd3c9b-83fe-44b6-a482-c2f3129dc4d9","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:08:40.604423Z","iopub.execute_input":"2025-07-24T07:08:40.604785Z","iopub.status.idle":"2025-07-24T07:08:49.520847Z","shell.execute_reply.started":"2025-07-24T07:08:40.604757Z","shell.execute_reply":"2025-07-24T07:08:49.515648Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### --- Generate submission file\nSort by ranker_id and the predicted score in descending order","metadata":{"_uuid":"08283ae6-f4d3-453a-8ed2-eee91a469235","_cell_guid":"2efc207b-38ea-4b29-accd-9760dadcc081","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"df_test_sorted = df_test_fe.sort_values(['ranker_id', 'y_pred'], ascending=[True, False])","metadata":{"_uuid":"ae579e23-da01-441c-b5bf-edc54ef04259","_cell_guid":"43f4091d-54c9-412f-b506-dc909549bcf2","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:09:21.030040Z","iopub.execute_input":"2025-07-24T07:09:21.030437Z","iopub.status.idle":"2025-07-24T07:09:33.386808Z","shell.execute_reply.started":"2025-07-24T07:09:21.030406Z","shell.execute_reply":"2025-07-24T07:09:33.383147Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Assign the rank (1, 2, 3, ...) for each flight within its search group","metadata":{"_uuid":"53256c7a-b2d0-4097-bbd3-36fc7723b177","_cell_guid":"94f8b2e4-d8a3-4d0e-8d3a-8ddae3832892","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"df_test_sorted['selected'] = df_test_sorted.groupby('ranker_id').cumcount() + 1","metadata":{"_uuid":"09f379cb-5422-421a-9479-94ccc402cdc5","_cell_guid":"6184d894-eb51-4122-a7bd-5e94f8c95eaa","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:09:38.522044Z","iopub.execute_input":"2025-07-24T07:09:38.522426Z","iopub.status.idle":"2025-07-24T07:09:40.326061Z","shell.execute_reply.started":"2025-07-24T07:09:38.522396Z","shell.execute_reply":"2025-07-24T07:09:40.322578Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Create the submission dataframe in the required format","metadata":{"_uuid":"925470b2-060b-4830-adb3-936b93081999","_cell_guid":"bb3477c7-ac3c-4805-919a-108b1a396623","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"submission = df_test_sorted[['Id', 'ranker_id', 'selected']]","metadata":{"_uuid":"1dfc4510-02ab-46c1-8d60-ff758ef58676","_cell_guid":"9d0f8426-7f22-47a0-b859-53d729350e27","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:09:42.413264Z","iopub.execute_input":"2025-07-24T07:09:42.413623Z","iopub.status.idle":"2025-07-24T07:09:42.654658Z","shell.execute_reply.started":"2025-07-24T07:09:42.413594Z","shell.execute_reply":"2025-07-24T07:09:42.648138Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Save to CSV","metadata":{"_uuid":"90a0fc54-ac74-4f94-82bf-f49036eb9f6e","_cell_guid":"0916f0b1-2651-46f9-a4f3-67383deaebe1","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index=False)\nprint(\"✅ Submission file 'submission.csv' created successfully!\")\nprint(submission.head())\nprint(submission.tail())","metadata":{"_uuid":"54e2958a-5f11-42a3-b147-d22d768df056","_cell_guid":"23d428fd-9fc3-4167-8b4a-8069aa472e94","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-07-24T07:09:44.673744Z","iopub.execute_input":"2025-07-24T07:09:44.674154Z","iopub.status.idle":"2025-07-24T07:09:57.385277Z","shell.execute_reply.started":"2025-07-24T07:09:44.674100Z","shell.execute_reply":"2025-07-24T07:09:57.381288Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null}]}