{
  "id": 590264,
  "title": "How was train.parquet generated from raw data?",
  "url": "/competitions/aeroclub-recsys-2025/discussion/590264",
  "author_name": "",
  "post_date": "2025-07-19T07:29:47.095522600Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi organizers,</p>\n<p>Thanks for providing both the raw data and the preprocessed file <code>train.parquet</code>.  <br>\nFor better understanding and to ensure consistency when I extract additional fields from the raw data,  <br>\ncould you share the code (or a simplified version) used to generate <code>train.parquet</code>?</p>\n<p>Having the preprocessing script would make it easier for participants to extend or reproduce the dataset  <br>\nwhile maintaining the same logic.</p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": "3250805",
      "postDate": "07/19/2025 07:29:47",
      "content": "<p>Hi organizers,</p>\n<p>Thanks for providing both the raw data and the preprocessed file <code>train.parquet</code>.  <br>\nFor better understanding and to ensure consistency when I extract additional fields from the raw data,  <br>\ncould you share the code (or a simplified version) used to generate <code>train.parquet</code>?</p>\n<p>Having the preprocessing script would make it easier for participants to extend or reproduce the dataset  <br>\nwhile maintaining the same logic.</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hi organizers,\n\nThanks for providing both the raw data and the preprocessed file `train.parquet`.  \nFor better understanding and to ensure consistency when I extract additional fields from the raw data,  \ncould you share the code (or a simplified version) used to generate `train.parquet`?\n\nHaving the preprocessing script would make it easier for participants to extend or reproduce the dataset  \nwhile maintaining the same logic.\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "3251008",
      "postDate": "07/19/2025 17:07:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a>. Thanks for asking. We've decided to share our <a href=\"https://www.kaggle.com/code/samvelkoch/reference-json-converters\" target=\"_blank\">json_converters</a> as a reference. Hope that helps. But please don't forget it couldn't be applied directly to the competition jsons since some ETL and sensitive data truncations. </p>",
      "rawMarkdown": "Hi @mango789. Thanks for asking. We've decided to share our [json_converters](https://www.kaggle.com/code/samvelkoch/reference-json-converters ) as a reference. Hope that helps. But please don't forget it couldn't be applied directly to the competition jsons since some ETL and sensitive data truncations.",
      "votes": null
    },
    {
      "id": "3251096",
      "postDate": "07/20/2025 00:01:12",
      "content": "<p>Thanks a lot for sharing the json_converters and the note about ETL/data truncations. That’s very helpful!🥺</p>",
      "rawMarkdown": "Thanks a lot for sharing the json_converters and the note about ETL/data truncations. That’s very helpful!🥺",
      "votes": null
    },
    {
      "id": "3253340",
      "postDate": "07/24/2025 13:45:56",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/samvelkoch\" target=\"_blank\">@samvelkoch</a> , I have a another quick question: within the JSON files, is the order of the flights preserved exactly in the parquet files for the same <code>ranker_id</code> group? In other words, can I assume the flights’ order in json matches the order in the parquet rows for that <code>ranker_id</code>?<br>\nThanks in advance for your help!</p>",
      "rawMarkdown": "Hi @samvelkoch , I have a another quick question: within the JSON files, is the order of the flights preserved exactly in the parquet files for the same `ranker_id` group? In other words, can I assume the flights’ order in json matches the order in the parquet rows for that `ranker_id`?\nThanks in advance for your help!",
      "votes": null
    },
    {
      "id": "3253358",
      "postDate": "07/24/2025 14:16:12",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a>. We just can't guarantee that. This assumption could not be 100% truth and I have no possibility to check it for sure now. So the answer I have to give you is NO. </p>",
      "rawMarkdown": "Hi @mango789. We just can't guarantee that. This assumption could not be 100% truth and I have no possibility to check it for sure now. So the answer I have to give you is NO.",
      "votes": null
    },
    {
      "id": "3253765",
      "postDate": "07/25/2025 08:09:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/samvelkoch\" target=\"_blank\">@samvelkoch</a>, thanks for the clarification.</p>\n<p>Is there any recommended way to reliably align the option order between the json file and the parquet file?<br>\nToday I tried joining them on the following keys:</p>\n<pre><code>on = [\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    \n]\n</code></pre>\n<p>But I still found about 300,000 rows couldn’t be matched correctly.<br>\nWould appreciate any advice on how to better align them. Thanks!</p>",
      "rawMarkdown": "Hi @samvelkoch, thanks for the clarification.\n\nIs there any recommended way to reliably align the option order between the json file and the parquet file?\nToday I tried joining them on the following keys:\n\n```python\non = [\n    \"ranker_id\",\n    \"legs0_departureAt\",\n    \"legs0_arrivalAt\",\n    \"legs1_departureAt\",\n    \"legs1_arrivalAt\",\n    \"legs0_segments0_flightNumber\",\n    \"legs1_segments0_flightNumber\",\n    \"totalPrice\"\n]\n```\n\nBut I still found about 300,000 rows couldn’t be matched correctly.\nWould appreciate any advice on how to better align them. Thanks!",
      "votes": null
    },
    {
      "id": "3253770",
      "postDate": "07/25/2025 08:15:07",
      "content": "<p>And I also tried grouping both the train and test datasets using:</p>\n<pre><code>train = pl.read_parquet(os.path.join(, )).drop()\ntest = (\n    pl.read_parquet(os.path.join(, ))\n    .drop()\n    .with_columns(pl.lit(, dtype=pl.Int64).alias())\n)\n\ntrain_test = pl.concat((train, test))\n\ntrain_test_df = train_test.group_by([\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n]).agg(pl.())\n\n(train_test_df)\n</code></pre>\n<p>This gives me a grouped dataframe of shape (8,165,696, 6).</p>\n<p>However, when I group the same data using the raw json files, I get only (7,489,609, 3) groups.</p>\n<p>Is there any known inconsistency or reason why the number of groups extracted from the parquet files differs significantly from the JSON source?</p>",
      "rawMarkdown": "And I also tried grouping both the train and test datasets using:\n\n```python\ntrain = pl.read_parquet(os.path.join(\"data\", \"train.parquet\")).drop(\"__index_level_0__\")\ntest = (\n    pl.read_parquet(os.path.join(\"data\", \"test.parquet\"))\n    .drop(\"__index_level_0__\")\n    .with_columns(pl.lit(0, dtype=pl.Int64).alias(\"selected\"))\n)\n\ntrain_test = pl.concat((train, test))\n\ntrain_test_df = train_test.group_by([\n    \"ranker_id\",\n    \"legs0_duration\",\n    \"legs1_duration\",\n    \"legs0_departureAt\",\n    \"legs0_arrivalAt\",\n    \"legs1_departureAt\",\n    \"legs1_arrivalAt\",\n    \"legs0_segments0_flightNumber\",\n    \"legs1_segments0_flightNumber\",\n]).agg(pl.len())\n\nprint(train_test_df)\n```\n\nThis gives me a grouped dataframe of shape (8,165,696, 6).\n\nHowever, when I group the same data using the raw json files, I get only (7,489,609, 3) groups.\n\nIs there any known inconsistency or reason why the number of groups extracted from the parquet files differs significantly from the JSON source?",
      "votes": null
    },
    {
      "id": "3253793",
      "postDate": "07/25/2025 09:16:34",
      "content": "<p>Thanks for diving so deep into the data <a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a>. I really would like to be helpful in both of your questions. We'll see if I found quick answer without hard reverse engineering research.  </p>",
      "rawMarkdown": "Thanks for diving so deep into the data @mango789. I really would like to be helpful in both of your questions. We'll see if I found quick answer without hard reverse engineering research.",
      "votes": null
    },
    {
      "id": "3253804",
      "postDate": "07/25/2025 09:41:54",
      "content": "<p>Thanks for the reply, <a href=\"https://www.kaggle.com/samvelkoch\" target=\"_blank\">@samvelkoch</a>. Appreciate you taking the time to look into it.</p>",
      "rawMarkdown": "Thanks for the reply, @samvelkoch. Appreciate you taking the time to look into it.",
      "votes": null
    },
    {
      "id": "3253916",
      "postDate": "07/25/2025 13:48:46",
      "content": "<p><a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a> it seems the issue roots are in the logic of data deduplication and cleaning in the ETL json -&gt; parquet pipeline. My suggestion is to focus on parquets as the primary source. It doesn't mean that jsons are useless. </p>",
      "rawMarkdown": "mango789 it seems the issue roots are in the logic of data deduplication and cleaning in the ETL json -> parquet pipeline. My suggestion is to focus on parquets as the primary source. It doesn't mean that jsons are useless.",
      "votes": null
    },
    {
      "id": "3269513",
      "postDate": "08/14/2025 15:22:01",
      "content": "<p>I also found some mismatch within these features from json to parquet file</p>\n<p>Also, I found the total data is different between json and parquet file although they have the same ranker_id.</p>\n<p>you can check with </p>\n<pre><code>, , , ,\n\n\ntarget_rid = \n\ntarget_rows = (\n    train_filled\n    .(pl.col() == target_rid)\n    .(pl.col().is_in([]))\n)\n\n\ntarget_rid = \n\ntarget_rows = (\n    train_filled\n    .(pl.col() == target_rid)\n    .(pl.col().is_in([, , ]))\n    .(pl.col() == )\n\n)\n\n\n\n\ntarget_rid = \n\ntarget_rows = (\n    train_filled\n    .(pl.col() == target_rid)\n    .(pl.col().is_in([, ]))\n    .(pl.col() == )\n\n)\n</code></pre>",
      "rawMarkdown": "I also found some mismatch within these features from json to parquet file\n\nAlso, I found the total data is different between json and parquet file although they have the same ranker_id.\n\n you can check with \n\n```python\n'legs0_segments1_baggageAllowance_quantity','legs0_segments1_seatsAvailable' , \"miniRules0_monetaryAmount\", \"miniRules1_monetaryAmount\",\n\n\ntarget_rid = \"0f4a3e1286794696a08882bbb7813497\"\n\ntarget_rows = (\n    train_filled\n    .filter(pl.col(\"ranker_id\") == target_rid)\n    .filter(pl.col(\"legs0_segments0_flightNumber\").is_in([\"208\"]))\n)\n\n\ntarget_rid = \"f7d878ee3ac944b38d084cf09e15f210\"\n\ntarget_rows = (\n    train_filled\n    .filter(pl.col(\"ranker_id\") == target_rid)#'legs0_segments2_baggageAllowance_quantity',\n    .filter(pl.col(\"legs0_segments1_flightNumber\").is_in([\"5309\", \"5305\", \"6371\"]))\n    .filter(pl.col(\"legs0_segments0_seatsAvailable\") == 2)\n\n)\n\n\n\n# 'legs0_segments2_seatsAvailable',#'legs0_segments2_baggageAllowance_quantity',\ntarget_rid = \"938b94a6d6584757a857ca180c9a060a\"\n\ntarget_rows = (\n    train_filled\n    .filter(pl.col(\"ranker_id\") == target_rid)\n    .filter(pl.col(\"legs0_segments0_flightNumber\").is_in([\"1008\", \"1004\"]))\n    .filter(pl.col(\"totalPrice\") == 262862.0)\n\n)\n\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3251008,
      "author_name": "samvelkoch",
      "author_url": "",
      "post_date": "07/19/2025 17:07:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a>. Thanks for asking. We've decided to share our <a href=\"https://www.kaggle.com/code/samvelkoch/reference-json-converters\" target=\"_blank\">json_converters</a> as a reference. Hope that helps. But please don't forget it couldn't be applied directly to the competition jsons since some ETL and sensitive data truncations. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3251096,
          "author_name": "mango789",
          "author_url": "",
          "post_date": "07/20/2025 00:01:12",
          "content": "<p>Thanks a lot for sharing the json_converters and the note about ETL/data truncations. That’s very helpful!🥺</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3253340,
          "author_name": "mango789",
          "author_url": "",
          "post_date": "07/24/2025 13:45:56",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/samvelkoch\" target=\"_blank\">@samvelkoch</a> , I have a another quick question: within the JSON files, is the order of the flights preserved exactly in the parquet files for the same <code>ranker_id</code> group? In other words, can I assume the flights’ order in json matches the order in the parquet rows for that <code>ranker_id</code>?<br>\nThanks in advance for your help!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3253358,
              "author_name": "samvelkoch",
              "author_url": "",
              "post_date": "07/24/2025 14:16:12",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a>. We just can't guarantee that. This assumption could not be 100% truth and I have no possibility to check it for sure now. So the answer I have to give you is NO. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3253765,
                  "author_name": "mango789",
                  "author_url": "",
                  "post_date": "07/25/2025 08:09:36",
                  "content": "<p>Hi <a href=\"https://www.kaggle.com/samvelkoch\" target=\"_blank\">@samvelkoch</a>, thanks for the clarification.</p>\n<p>Is there any recommended way to reliably align the option order between the json file and the parquet file?<br>\nToday I tried joining them on the following keys:</p>\n<pre><code>on = [\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    \n]\n</code></pre>\n<p>But I still found about 300,000 rows couldn’t be matched correctly.<br>\nWould appreciate any advice on how to better align them. Thanks!</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3253770,
                  "author_name": "mango789",
                  "author_url": "",
                  "post_date": "07/25/2025 08:15:07",
                  "content": "<p>And I also tried grouping both the train and test datasets using:</p>\n<pre><code>train = pl.read_parquet(os.path.join(, )).drop()\ntest = (\n    pl.read_parquet(os.path.join(, ))\n    .drop()\n    .with_columns(pl.lit(, dtype=pl.Int64).alias())\n)\n\ntrain_test = pl.concat((train, test))\n\ntrain_test_df = train_test.group_by([\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n    ,\n]).agg(pl.())\n\n(train_test_df)\n</code></pre>\n<p>This gives me a grouped dataframe of shape (8,165,696, 6).</p>\n<p>However, when I group the same data using the raw json files, I get only (7,489,609, 3) groups.</p>\n<p>Is there any known inconsistency or reason why the number of groups extracted from the parquet files differs significantly from the JSON source?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3253793,
                      "author_name": "samvelkoch",
                      "author_url": "",
                      "post_date": "07/25/2025 09:16:34",
                      "content": "<p>Thanks for diving so deep into the data <a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a>. I really would like to be helpful in both of your questions. We'll see if I found quick answer without hard reverse engineering research.  </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3253804,
                          "author_name": "mango789",
                          "author_url": "",
                          "post_date": "07/25/2025 09:41:54",
                          "content": "<p>Thanks for the reply, <a href=\"https://www.kaggle.com/samvelkoch\" target=\"_blank\">@samvelkoch</a>. Appreciate you taking the time to look into it.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3253916,
                              "author_name": "samvelkoch",
                              "author_url": "",
                              "post_date": "07/25/2025 13:48:46",
                              "content": "<p><a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a> it seems the issue roots are in the logic of data deduplication and cleaning in the ETL json -&gt; parquet pipeline. My suggestion is to focus on parquets as the primary source. It doesn't mean that jsons are useless. </p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3269513,
      "author_name": "dingyangwang",
      "author_url": "",
      "post_date": "08/14/2025 15:22:01",
      "content": "<p>I also found some mismatch within these features from json to parquet file</p>\n<p>Also, I found the total data is different between json and parquet file although they have the same ranker_id.</p>\n<p>you can check with </p>\n<pre><code>, , , ,\n\n\ntarget_rid = \n\ntarget_rows = (\n    train_filled\n    .(pl.col() == target_rid)\n    .(pl.col().is_in([]))\n)\n\n\ntarget_rid = \n\ntarget_rows = (\n    train_filled\n    .(pl.col() == target_rid)\n    .(pl.col().is_in([, , ]))\n    .(pl.col() == )\n\n)\n\n\n\n\ntarget_rid = \n\ntarget_rows = (\n    train_filled\n    .(pl.col() == target_rid)\n    .(pl.col().is_in([, ]))\n    .(pl.col() == )\n\n)\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3250805": "Hi organizers,\n\nThanks for providing both the raw data and the preprocessed file `train.parquet`.  \nFor better understanding and to ensure consistency when I extract additional fields from the raw data,  \ncould you share the code (or a simplified version) used to generate `train.parquet`?\n\nHaving the preprocessing script would make it easier for participants to extend or reproduce the dataset  \nwhile maintaining the same logic.\n\nThanks in advance!",
    "3251008": "Hi @mango789. Thanks for asking. We've decided to share our [json_converters](https://www.kaggle.com/code/samvelkoch/reference-json-converters ) as a reference. Hope that helps. But please don't forget it couldn't be applied directly to the competition jsons since some ETL and sensitive data truncations.",
    "3251096": "Thanks a lot for sharing the json_converters and the note about ETL/data truncations. That’s very helpful!🥺",
    "3253340": "Hi @samvelkoch , I have a another quick question: within the JSON files, is the order of the flights preserved exactly in the parquet files for the same `ranker_id` group? In other words, can I assume the flights’ order in json matches the order in the parquet rows for that `ranker_id`?\nThanks in advance for your help!",
    "3253358": "Hi @mango789. We just can't guarantee that. This assumption could not be 100% truth and I have no possibility to check it for sure now. So the answer I have to give you is NO.",
    "3253765": "Hi @samvelkoch, thanks for the clarification.\n\nIs there any recommended way to reliably align the option order between the json file and the parquet file?\nToday I tried joining them on the following keys:\n\n```python\non = [\n    \"ranker_id\",\n    \"legs0_departureAt\",\n    \"legs0_arrivalAt\",\n    \"legs1_departureAt\",\n    \"legs1_arrivalAt\",\n    \"legs0_segments0_flightNumber\",\n    \"legs1_segments0_flightNumber\",\n    \"totalPrice\"\n]\n```\n\nBut I still found about 300,000 rows couldn’t be matched correctly.\nWould appreciate any advice on how to better align them. Thanks!",
    "3253770": "And I also tried grouping both the train and test datasets using:\n\n```python\ntrain = pl.read_parquet(os.path.join(\"data\", \"train.parquet\")).drop(\"__index_level_0__\")\ntest = (\n    pl.read_parquet(os.path.join(\"data\", \"test.parquet\"))\n    .drop(\"__index_level_0__\")\n    .with_columns(pl.lit(0, dtype=pl.Int64).alias(\"selected\"))\n)\n\ntrain_test = pl.concat((train, test))\n\ntrain_test_df = train_test.group_by([\n    \"ranker_id\",\n    \"legs0_duration\",\n    \"legs1_duration\",\n    \"legs0_departureAt\",\n    \"legs0_arrivalAt\",\n    \"legs1_departureAt\",\n    \"legs1_arrivalAt\",\n    \"legs0_segments0_flightNumber\",\n    \"legs1_segments0_flightNumber\",\n]).agg(pl.len())\n\nprint(train_test_df)\n```\n\nThis gives me a grouped dataframe of shape (8,165,696, 6).\n\nHowever, when I group the same data using the raw json files, I get only (7,489,609, 3) groups.\n\nIs there any known inconsistency or reason why the number of groups extracted from the parquet files differs significantly from the JSON source?",
    "3253793": "Thanks for diving so deep into the data @mango789. I really would like to be helpful in both of your questions. We'll see if I found quick answer without hard reverse engineering research.",
    "3253804": "Thanks for the reply, @samvelkoch. Appreciate you taking the time to look into it.",
    "3253916": "mango789 it seems the issue roots are in the logic of data deduplication and cleaning in the ETL json -> parquet pipeline. My suggestion is to focus on parquets as the primary source. It doesn't mean that jsons are useless.",
    "3269513": "I also found some mismatch within these features from json to parquet file\n\nAlso, I found the total data is different between json and parquet file although they have the same ranker_id.\n\n you can check with \n\n```python\n'legs0_segments1_baggageAllowance_quantity','legs0_segments1_seatsAvailable' , \"miniRules0_monetaryAmount\", \"miniRules1_monetaryAmount\",\n\n\ntarget_rid = \"0f4a3e1286794696a08882bbb7813497\"\n\ntarget_rows = (\n    train_filled\n    .filter(pl.col(\"ranker_id\") == target_rid)\n    .filter(pl.col(\"legs0_segments0_flightNumber\").is_in([\"208\"]))\n)\n\n\ntarget_rid = \"f7d878ee3ac944b38d084cf09e15f210\"\n\ntarget_rows = (\n    train_filled\n    .filter(pl.col(\"ranker_id\") == target_rid)#'legs0_segments2_baggageAllowance_quantity',\n    .filter(pl.col(\"legs0_segments1_flightNumber\").is_in([\"5309\", \"5305\", \"6371\"]))\n    .filter(pl.col(\"legs0_segments0_seatsAvailable\") == 2)\n\n)\n\n\n\n# 'legs0_segments2_seatsAvailable',#'legs0_segments2_baggageAllowance_quantity',\ntarget_rid = \"938b94a6d6584757a857ca180c9a060a\"\n\ntarget_rows = (\n    train_filled\n    .filter(pl.col(\"ranker_id\") == target_rid)\n    .filter(pl.col(\"legs0_segments0_flightNumber\").is_in([\"1008\", \"1004\"]))\n    .filter(pl.col(\"totalPrice\") == 262862.0)\n\n)\n\n```"
  },
  "source": "meta"
}