{
  "id": 583971,
  "title": "How to make sense of the prediction problem given the limited information? ",
  "url": "/competitions/social-sim-challenge-social-media-based-personas/discussion/583971",
  "author_name": "",
  "post_date": "2025-06-10T18:48:48.880488Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, first, thank you to the organizers for hosting this competition!</p>\n<p>From a quick glance at the data, we observe the following:</p>\n<p>1) Each user ID appears only once or twice across the dataset, and there is no overlap of users between the train, validation, and test sets.</p>\n<p>2) This means we are essentially predicting a user’s action based solely on the post they are interacting with and the timestamp. Given that we have no user-specific information, is it reasonable to expect a model to accurately distinguish between follow, like, or dislike from just this limited context?</p>",
  "messages": [
    {
      "id": "3221320",
      "postDate": "06/10/2025 18:48:48",
      "content": "<p>Hi, first, thank you to the organizers for hosting this competition!</p>\n<p>From a quick glance at the data, we observe the following:</p>\n<p>1) Each user ID appears only once or twice across the dataset, and there is no overlap of users between the train, validation, and test sets.</p>\n<p>2) This means we are essentially predicting a user’s action based solely on the post they are interacting with and the timestamp. Given that we have no user-specific information, is it reasonable to expect a model to accurately distinguish between follow, like, or dislike from just this limited context?</p>",
      "rawMarkdown": "Hi, first, thank you to the organizers for hosting this competition!\n\nFrom a quick glance at the data, we observe the following:\n\n1) Each user ID appears only once or twice across the dataset, and there is no overlap of users between the train, validation, and test sets.\n\n2) This means we are essentially predicting a user’s action based solely on the post they are interacting with and the timestamp. Given that we have no user-specific information, is it reasonable to expect a model to accurately distinguish between follow, like, or dislike from just this limited context?",
      "votes": null
    },
    {
      "id": "3222070",
      "postDate": "06/11/2025 19:05:13",
      "content": "<p>Hello, the <code>user_id</code> field is mainly used to distinguish users between the threads. The last user belongs to the \"cluster\". The task is to predict on the cluster level the right action given the post. Since it is on the cluster level, we do not expect the model to provide 90%+ accuracy. </p>",
      "rawMarkdown": "Hello, the `user_id` field is mainly used to distinguish users between the threads. The last user belongs to the \"cluster\". The task is to predict on the cluster level the right action given the post. Since it is on the cluster level, we do not expect the model to provide 90%+ accuracy.",
      "votes": null
    },
    {
      "id": "3222226",
      "postDate": "06/12/2025 03:32:59",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3222070,
      "author_name": "rstzzzz",
      "author_url": "",
      "post_date": "06/11/2025 19:05:13",
      "content": "<p>Hello, the <code>user_id</code> field is mainly used to distinguish users between the threads. The last user belongs to the \"cluster\". The task is to predict on the cluster level the right action given the post. Since it is on the cluster level, we do not expect the model to provide 90%+ accuracy. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3222226,
      "author_name": "tianyipeng",
      "author_url": "",
      "post_date": "06/12/2025 03:32:59",
      "content": "<p>Thank you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3221320": "Hi, first, thank you to the organizers for hosting this competition!\n\nFrom a quick glance at the data, we observe the following:\n\n1) Each user ID appears only once or twice across the dataset, and there is no overlap of users between the train, validation, and test sets.\n\n2) This means we are essentially predicting a user’s action based solely on the post they are interacting with and the timestamp. Given that we have no user-specific information, is it reasonable to expect a model to accurately distinguish between follow, like, or dislike from just this limited context?",
    "3222070": "Hello, the `user_id` field is mainly used to distinguish users between the threads. The last user belongs to the \"cluster\". The task is to predict on the cluster level the right action given the post. Since it is on the cluster level, we do not expect the model to provide 90%+ accuracy.",
    "3222226": "Thank you!"
  },
  "source": "meta"
}