{
  "id": 586699,
  "title": "Curious observation: Are dataset rows in the same order as users saw them?",
  "url": "/competitions/aeroclub-recsys-2025/discussion/586699",
  "author_name": "",
  "post_date": "2025-06-27T18:40:11.086831400Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>While exploring the dataset, I noticed that the rows don't appear to be shuffled. Interestingly, the first candidate in each set is selected noticeably more often than others. Could it be that the options are listed in the same order users originally saw them? That might explain the higher selection rate for the first position.</p>\n<p>To check this, I plotted two distributions (on a log-scale x-axis):</p>\n<ul>\n<li>All relative positions of candidates within <code>ranker_id</code></li>\n<li>Relative positions of <strong>selected</strong> candidates</li>\n</ul>\n<p>The selected ones are heavily skewed toward the first few positions. I also applied the same analysis on the test set using my model’s predictions (HitRate@3 = 0.49) — and saw a very similar skew among predicted selections. Anyone looked into this further? Maybe raw JSON data has clues?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3494315%2F2c39d7dbed667f91e3a4b8a984fac2ea%2Fraw-corr.png?generation=1751049239445756&amp;alt=media\" alt=\"Distributions\"></p>",
  "messages": [
    {
      "id": "3234381",
      "postDate": "06/27/2025 18:40:11",
      "content": "<p>While exploring the dataset, I noticed that the rows don't appear to be shuffled. Interestingly, the first candidate in each set is selected noticeably more often than others. Could it be that the options are listed in the same order users originally saw them? That might explain the higher selection rate for the first position.</p>\n<p>To check this, I plotted two distributions (on a log-scale x-axis):</p>\n<ul>\n<li>All relative positions of candidates within <code>ranker_id</code></li>\n<li>Relative positions of <strong>selected</strong> candidates</li>\n</ul>\n<p>The selected ones are heavily skewed toward the first few positions. I also applied the same analysis on the test set using my model’s predictions (HitRate@3 = 0.49) — and saw a very similar skew among predicted selections. Anyone looked into this further? Maybe raw JSON data has clues?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3494315%2F2c39d7dbed667f91e3a4b8a984fac2ea%2Fraw-corr.png?generation=1751049239445756&amp;alt=media\" alt=\"Distributions\"></p>",
      "rawMarkdown": "While exploring the dataset, I noticed that the rows don't appear to be shuffled. Interestingly, the first candidate in each set is selected noticeably more often than others. Could it be that the options are listed in the same order users originally saw them? That might explain the higher selection rate for the first position.\n\nTo check this, I plotted two distributions (on a log-scale x-axis):\n- All relative positions of candidates within `ranker_id`\n- Relative positions of **selected** candidates\n\nThe selected ones are heavily skewed toward the first few positions. I also applied the same analysis on the test set using my model’s predictions (HitRate@3 = 0.49) — and saw a very similar skew among predicted selections. Anyone looked into this further? Maybe raw JSON data has clues?\n\n![Distributions](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3494315%2F2c39d7dbed667f91e3a4b8a984fac2ea%2Fraw-corr.png?generation=1751049239445756&alt=media)",
      "votes": null
    },
    {
      "id": "3235072",
      "postDate": "06/28/2025 17:15:45",
      "content": "<p>Interesting insight <a href=\"https://www.kaggle.com/ka1242\" target=\"_blank\">@ka1242</a> 👏 Thank you for sharing this promptly. One of the reasons why the Kaggle community is so effective at solving problems.</p>\n<p>I can confirm that there is indeed a positional shift in the data - this is a characteristic of how search results are formed at a low level. The shift is due to how flights options assembled from different suppliers and airlines. The process is non-deterministic, so the results can differ each time.</p>\n<p><strong>We strongly recommend that competition participants do NOT use positional features as an exploit and do not base their models on this knowledge.</strong> Models with such features are typically useless. Features derived from positional biases in training data are HARMFUL, unstable, and as our experiments have shown, create incorrect bias. </p>\n<p>Additionally, the competition metric that doesn't consider ranks for groups with 10 or fewer elements significantly diminishes the impact of such positional features on the score.</p>",
      "rawMarkdown": "Interesting insight @ka1242 👏 Thank you for sharing this promptly. One of the reasons why the Kaggle community is so effective at solving problems.\n\nI can confirm that there is indeed a positional shift in the data - this is a characteristic of how search results are formed at a low level. The shift is due to how flights options assembled from different suppliers and airlines. The process is non-deterministic, so the results can differ each time.\n\n**We strongly recommend that competition participants do NOT use positional features as an exploit and do not base their models on this knowledge.** Models with such features are typically useless. Features derived from positional biases in training data are HARMFUL, unstable, and as our experiments have shown, create incorrect bias. \n\nAdditionally, the competition metric that doesn't consider ranks for groups with 10 or fewer elements significantly diminishes the impact of such positional features on the score.",
      "votes": null
    },
    {
      "id": "3235189",
      "postDate": "06/28/2025 20:36:03",
      "content": "<p>As far as I remember from the time I often fly system normally shows direct flies first. So if there is only 1 -2 direct flights between the cities, they probably will be first in list and that what user probably will pick up.  May be this explains….</p>",
      "rawMarkdown": "As far as I remember from the time I often fly system normally shows direct flies first. So if there is only 1 -2 direct flights between the cities, they probably will be first in list and that what user probably will pick up.  May be this explains....",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3235072,
      "author_name": "samvelkoch",
      "author_url": "",
      "post_date": "06/28/2025 17:15:45",
      "content": "<p>Interesting insight <a href=\"https://www.kaggle.com/ka1242\" target=\"_blank\">@ka1242</a> 👏 Thank you for sharing this promptly. One of the reasons why the Kaggle community is so effective at solving problems.</p>\n<p>I can confirm that there is indeed a positional shift in the data - this is a characteristic of how search results are formed at a low level. The shift is due to how flights options assembled from different suppliers and airlines. The process is non-deterministic, so the results can differ each time.</p>\n<p><strong>We strongly recommend that competition participants do NOT use positional features as an exploit and do not base their models on this knowledge.</strong> Models with such features are typically useless. Features derived from positional biases in training data are HARMFUL, unstable, and as our experiments have shown, create incorrect bias. </p>\n<p>Additionally, the competition metric that doesn't consider ranks for groups with 10 or fewer elements significantly diminishes the impact of such positional features on the score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3235189,
      "author_name": "sergeyqt2024",
      "author_url": "",
      "post_date": "06/28/2025 20:36:03",
      "content": "<p>As far as I remember from the time I often fly system normally shows direct flies first. So if there is only 1 -2 direct flights between the cities, they probably will be first in list and that what user probably will pick up.  May be this explains….</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3234381": "While exploring the dataset, I noticed that the rows don't appear to be shuffled. Interestingly, the first candidate in each set is selected noticeably more often than others. Could it be that the options are listed in the same order users originally saw them? That might explain the higher selection rate for the first position.\n\nTo check this, I plotted two distributions (on a log-scale x-axis):\n- All relative positions of candidates within `ranker_id`\n- Relative positions of **selected** candidates\n\nThe selected ones are heavily skewed toward the first few positions. I also applied the same analysis on the test set using my model’s predictions (HitRate@3 = 0.49) — and saw a very similar skew among predicted selections. Anyone looked into this further? Maybe raw JSON data has clues?\n\n![Distributions](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3494315%2F2c39d7dbed667f91e3a4b8a984fac2ea%2Fraw-corr.png?generation=1751049239445756&alt=media)",
    "3235072": "Interesting insight @ka1242 👏 Thank you for sharing this promptly. One of the reasons why the Kaggle community is so effective at solving problems.\n\nI can confirm that there is indeed a positional shift in the data - this is a characteristic of how search results are formed at a low level. The shift is due to how flights options assembled from different suppliers and airlines. The process is non-deterministic, so the results can differ each time.\n\n**We strongly recommend that competition participants do NOT use positional features as an exploit and do not base their models on this knowledge.** Models with such features are typically useless. Features derived from positional biases in training data are HARMFUL, unstable, and as our experiments have shown, create incorrect bias. \n\nAdditionally, the competition metric that doesn't consider ranks for groups with 10 or fewer elements significantly diminishes the impact of such positional features on the score.",
    "3235189": "As far as I remember from the time I often fly system normally shows direct flies first. So if there is only 1 -2 direct flights between the cities, they probably will be first in list and that what user probably will pick up.  May be this explains...."
  },
  "source": "meta"
}