{
  "id": 599563,
  "title": "Reflections after competition – real data ≠ easy scores",
  "url": "/competitions/aeroclub-recsys-2025/discussion/599563",
  "author_name": "",
  "post_date": "2025-08-17T11:28:57.822173300Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>I once again realized something:<br>\non real-world data, getting a strong leaderboard score is way harder than just improving CV.</p>\n<p>In toy datasets, you can often push the score higher with a better model, some feature crosses, or fine-tuned validation.<br>\nBut here – with real business traveler choices – that wasn’t enough.</p>\n<p>It made me think: how should we actually approach score improvement in such competitions?</p>\n<p>Should we look for external knowledge (e.g. business routes, corporate travel policies)?</p>\n<p>Should we analyze the hardest cases (sessions where the model’s prediction is completely off)?</p>\n<p>Or does it come down to resources and prize pool size – bigger prizes bring more teams and stronger top solutions?</p>\n<p>I feel like in tasks like this, domain understanding is just as important as pure ML tricks.</p>\n<p>What do you think? How do you approach these “messy real-world” datasets?</p>\n<p>And a question for the organizers: 👉 will there be another competition in this space?<br>\nI’d love to see more challenges like it.</p>",
  "messages": [
    {
      "id": "3270808",
      "postDate": "08/17/2025 11:28:57",
      "content": "<p>Hey everyone,</p>\n<p>I once again realized something:<br>\non real-world data, getting a strong leaderboard score is way harder than just improving CV.</p>\n<p>In toy datasets, you can often push the score higher with a better model, some feature crosses, or fine-tuned validation.<br>\nBut here – with real business traveler choices – that wasn’t enough.</p>\n<p>It made me think: how should we actually approach score improvement in such competitions?</p>\n<p>Should we look for external knowledge (e.g. business routes, corporate travel policies)?</p>\n<p>Should we analyze the hardest cases (sessions where the model’s prediction is completely off)?</p>\n<p>Or does it come down to resources and prize pool size – bigger prizes bring more teams and stronger top solutions?</p>\n<p>I feel like in tasks like this, domain understanding is just as important as pure ML tricks.</p>\n<p>What do you think? How do you approach these “messy real-world” datasets?</p>\n<p>And a question for the organizers: 👉 will there be another competition in this space?<br>\nI’d love to see more challenges like it.</p>",
      "rawMarkdown": "Hey everyone,\n\nI once again realized something:\non real-world data, getting a strong leaderboard score is way harder than just improving CV.\n\nIn toy datasets, you can often push the score higher with a better model, some feature crosses, or fine-tuned validation.\nBut here – with real business traveler choices – that wasn’t enough.\n\nIt made me think: how should we actually approach score improvement in such competitions?\n\nShould we look for external knowledge (e.g. business routes, corporate travel policies)?\n\nShould we analyze the hardest cases (sessions where the model’s prediction is completely off)?\n\nOr does it come down to resources and prize pool size – bigger prizes bring more teams and stronger top solutions?\n\nI feel like in tasks like this, domain understanding is just as important as pure ML tricks.\n\nWhat do you think? How do you approach these “messy real-world” datasets?\n\nAnd a question for the organizers: 👉 will there be another competition in this space?\nI’d love to see more challenges like it.",
      "votes": null
    },
    {
      "id": "3270819",
      "postDate": "08/17/2025 12:16:38",
      "content": "<p>Thanks for this topic <a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a>. Yes there is  tons of data and ways to research more. Clickstream data, multimodal trips with different services, travel policies, market trends, agents/clients communications, etc.</p>\n<p>Looking forward for more challenges soon… </p>",
      "rawMarkdown": "Thanks for this topic @qurusx. Yes there is  tons of data and ways to research more. Clickstream data, multimodal trips with different services, travel policies, market trends, agents/clients communications, etc.\n\nLooking forward for more challenges soon...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3270819,
      "author_name": "samvelkoch",
      "author_url": "",
      "post_date": "08/17/2025 12:16:38",
      "content": "<p>Thanks for this topic <a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a>. Yes there is  tons of data and ways to research more. Clickstream data, multimodal trips with different services, travel policies, market trends, agents/clients communications, etc.</p>\n<p>Looking forward for more challenges soon… </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3270808": "Hey everyone,\n\nI once again realized something:\non real-world data, getting a strong leaderboard score is way harder than just improving CV.\n\nIn toy datasets, you can often push the score higher with a better model, some feature crosses, or fine-tuned validation.\nBut here – with real business traveler choices – that wasn’t enough.\n\nIt made me think: how should we actually approach score improvement in such competitions?\n\nShould we look for external knowledge (e.g. business routes, corporate travel policies)?\n\nShould we analyze the hardest cases (sessions where the model’s prediction is completely off)?\n\nOr does it come down to resources and prize pool size – bigger prizes bring more teams and stronger top solutions?\n\nI feel like in tasks like this, domain understanding is just as important as pure ML tricks.\n\nWhat do you think? How do you approach these “messy real-world” datasets?\n\nAnd a question for the organizers: 👉 will there be another competition in this space?\nI’d love to see more challenges like it.",
    "3270819": "Thanks for this topic @qurusx. Yes there is  tons of data and ways to research more. Clickstream data, multimodal trips with different services, travel policies, market trends, agents/clients communications, etc.\n\nLooking forward for more challenges soon..."
  },
  "source": "meta"
}