{
  "id": 41650,
  "title": "Product ID and category_id are related.",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/41650",
  "author_name": "",
  "post_date": "2017-10-22T00:17:56.612227200Z",
  "votes": 9,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I noticed that product id values have strong relation to category_id. \nHere's the histogram of top 10 frequent items by product id (X-axis label 0.5 means product id 5,000,000, 2.0 is 20,000,000, etc)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/233979/7722/info_leak.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>From the histogram, you can find Category_id 1000010653 is very frequent in low product _id values and not much in 100k to 1,000k range. </p>\n\n<p>Someone will figure out this eventually(or some are using it already) and use this information, so I guess it is better to disclose early and get some feedback from Kaggle admin. </p>\n\n<p>Using this information is not beneficial for both competitors and competition host since it does not give any educational value and commercial value. So I hope that the administrator bans use of this information - at least for prize winner (it is practically impossible to ban for others, so it is up to each participants conscience anyway.  :) )</p>\n\n<p>BTW, if you are splitting the train set into several records by product id sequentially, you'd better shuffle it first.</p>",
  "messages": [
    {
      "id": "233979",
      "postDate": "10/22/2017 00:17:56",
      "content": "<p>I noticed that product id values have strong relation to category_id. \nHere's the histogram of top 10 frequent items by product id (X-axis label 0.5 means product id 5,000,000, 2.0 is 20,000,000, etc)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/233979/7722/info_leak.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>From the histogram, you can find Category_id 1000010653 is very frequent in low product _id values and not much in 100k to 1,000k range. </p>\n\n<p>Someone will figure out this eventually(or some are using it already) and use this information, so I guess it is better to disclose early and get some feedback from Kaggle admin. </p>\n\n<p>Using this information is not beneficial for both competitors and competition host since it does not give any educational value and commercial value. So I hope that the administrator bans use of this information - at least for prize winner (it is practically impossible to ban for others, so it is up to each participants conscience anyway.  :) )</p>\n\n<p>BTW, if you are splitting the train set into several records by product id sequentially, you'd better shuffle it first.</p>",
      "rawMarkdown": "I noticed that product id values have strong relation to category_id. \nHere's the histogram of top 10 frequent items by product id (X-axis label 0.5 means product id 5,000,000, 2.0 is 20,000,000, etc)\n\n![enter image description here][1]\n\nFrom the histogram, you can find Category_id 1000010653 is very frequent in low product _id values and not much in 100k to 1,000k range. \n\nSomeone will figure out this eventually(or some are using it already) and use this information, so I guess it is better to disclose early and get some feedback from Kaggle admin. \n\nUsing this information is not beneficial for both competitors and competition host since it does not give any educational value and commercial value. So I hope that the administrator bans use of this information - at least for prize winner (it is practically impossible to ban for others, so it is up to each participants conscience anyway.  :) )\n\nBTW, if you are splitting the train set into several records by product id sequentially, you'd better shuffle it first.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/233979/7722/info_leak.png",
      "votes": null
    },
    {
      "id": "233980",
      "postDate": "10/22/2017 00:22:33",
      "content": "<p>Code for above histogram is:</p>\n\n<pre><code># Assume train_df is dataframe of _id and category_id\nfreq_df = pd.crosstab(index=train_df['category_id'], columns='count').sort_values('count', axis=0, ascending=False)\ntop10s = dict([(i, train_df[df['category_id'] == i].index) for i in freq_df.index[:10]])\ndf = pd.DataFrame.from_dict(top10s, orient='index').T\ndf.plot.hist(stacked=True, bins=100, figsize=(16, 6))\n</code></pre>",
      "rawMarkdown": "Code for above histogram is:\n\n    # Assume train_df is dataframe of _id and category_id\n    freq_df = pd.crosstab(index=train_df['category_id'], columns='count').sort_values('count', axis=0, ascending=False)\n    top10s = dict([(i, train_df[df['category_id'] == i].index) for i in freq_df.index[:10]])\n    df = pd.DataFrame.from_dict(top10s, orient='index').T\n    df.plot.hist(stacked=True, bins=100, figsize=(16, 6))",
      "votes": null
    },
    {
      "id": "234025",
      "postDate": "10/22/2017 03:51:18",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "234027",
      "postDate": "10/22/2017 03:51:38",
      "content": "<p>You are noticing some serious problem, that could be utilized as cheating. Luckily we do split it randomly. I am doubting if the undisclosed data are as same pattern as your histogram. Did you test that?</p>",
      "rawMarkdown": "You are noticing some serious problem, that could be utilized as cheating. Luckily we do split it randomly. I am doubting if the undisclosed data are as same pattern as your histogram. Did you test that?",
      "votes": null
    },
    {
      "id": "234031",
      "postDate": "10/22/2017 04:03:46",
      "content": "<p>You can see some missing _id in train set, and they are in test set. I believe that test set's _id also follows same pattern as train set's _id.</p>",
      "rawMarkdown": "You can see some missing _id in train set, and they are in test set. I believe that test set's _id also follows same pattern as train set's _id.",
      "votes": null
    },
    {
      "id": "234038",
      "postDate": "10/22/2017 04:54:56",
      "content": "<p>&gt;&gt; missing _id in train set, and they are in test set</p>\n\n<p>does it also means that if train_product_id is close to test_product_id, their label would be close? It may not be only for Category_id 1000010653?</p>",
      "rawMarkdown": "&gt;&gt; missing _id in train set, and they are in test set\n\ndoes it also means that if train_product_id is close to test_product_id, their label would be close? It may not be only for Category_id 1000010653?",
      "votes": null
    },
    {
      "id": "234039",
      "postDate": "10/22/2017 04:56:58",
      "content": "<p>From your chart, actually, only 1/10, say,  <code>product_id 10000010653</code> have the problem you described, others are a uniform distribution. So it that really a serious problem?</p>",
      "rawMarkdown": "From your chart, actually, only 1/10, say,  `product_id 10000010653` have the problem you described, others are a uniform distribution. So it that really a serious problem?",
      "votes": null
    },
    {
      "id": "234040",
      "postDate": "10/22/2017 04:59:44",
      "content": "<p>I think the answer is not. Actually, in the training set, close <code>product_id</code> doesn't have close <code>category_id</code>. Why do you think so?</p>",
      "rawMarkdown": "I think the answer is not. Actually, in the training set, close `product_id` doesn't have close `category_id`. Why do you think so?",
      "votes": null
    },
    {
      "id": "234063",
      "postDate": "10/22/2017 07:32:24",
      "content": "<p>I can say brown category doesn't occur at _id between 2e6 and 6e6 and purple one doesn't occur between 3e6 and 6e6, and this pattern applies to 8 of 10 frequent items. I applied this info to my local validation result, but I see improvements on only really tiny samples(0.0001%).  </p>\n\n<p>Fortunately, I believe that this will not be much an issue for this competition :) </p>\n\n<p>But it would be good to hear some feedback from the host or admin.</p>",
      "rawMarkdown": "I can say brown category doesn't occur at _id between 2e6 and 6e6 and purple one doesn't occur between 3e6 and 6e6, and this pattern applies to 8 of 10 frequent items. I applied this info to my local validation result, but I see improvements on only really tiny samples(0.0001%).  \n\nFortunately, I believe that this will not be much an issue for this competition :) \n\nBut it would be good to hear some feedback from the host or admin.",
      "votes": null
    },
    {
      "id": "234064",
      "postDate": "10/22/2017 07:35:57",
      "content": "<p>@Heng \nYes. Probability of having same label is higher if train/test product _id value is close.</p>",
      "rawMarkdown": "Heng \nYes. Probability of having same label is higher if train/test product _id value is close.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 233980,
      "author_name": "jandjenter",
      "author_url": "",
      "post_date": "10/22/2017 00:22:33",
      "content": "<p>Code for above histogram is:</p>\n\n<pre><code># Assume train_df is dataframe of _id and category_id\nfreq_df = pd.crosstab(index=train_df['category_id'], columns='count').sort_values('count', axis=0, ascending=False)\ntop10s = dict([(i, train_df[df['category_id'] == i].index) for i in freq_df.index[:10]])\ndf = pd.DataFrame.from_dict(top10s, orient='index').T\ndf.plot.hist(stacked=True, bins=100, figsize=(16, 6))\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 234027,
          "author_name": "heyheyforhello",
          "author_url": "",
          "post_date": "10/22/2017 03:51:38",
          "content": "<p>You are noticing some serious problem, that could be utilized as cheating. Luckily we do split it randomly. I am doubting if the undisclosed data are as same pattern as your histogram. Did you test that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234031,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "10/22/2017 04:03:46",
          "content": "<p>You can see some missing _id in train set, and they are in test set. I believe that test set's _id also follows same pattern as train set's _id.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234038,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/22/2017 04:54:56",
          "content": "<p>&gt;&gt; missing _id in train set, and they are in test set</p>\n\n<p>does it also means that if train_product_id is close to test_product_id, their label would be close? It may not be only for Category_id 1000010653?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234040,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/22/2017 04:59:44",
          "content": "<p>I think the answer is not. Actually, in the training set, close <code>product_id</code> doesn't have close <code>category_id</code>. Why do you think so?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 234064,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "10/22/2017 07:35:57",
          "content": "<p>@Heng \nYes. Probability of having same label is higher if train/test product _id value is close.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 234025,
      "author_name": "heyheyforhello",
      "author_url": "",
      "post_date": "10/22/2017 03:51:18",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 234039,
      "author_name": "brianlzm",
      "author_url": "",
      "post_date": "10/22/2017 04:56:58",
      "content": "<p>From your chart, actually, only 1/10, say,  <code>product_id 10000010653</code> have the problem you described, others are a uniform distribution. So it that really a serious problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 234063,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "10/22/2017 07:32:24",
          "content": "<p>I can say brown category doesn't occur at _id between 2e6 and 6e6 and purple one doesn't occur between 3e6 and 6e6, and this pattern applies to 8 of 10 frequent items. I applied this info to my local validation result, but I see improvements on only really tiny samples(0.0001%).  </p>\n\n<p>Fortunately, I believe that this will not be much an issue for this competition :) </p>\n\n<p>But it would be good to hear some feedback from the host or admin.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "233979": "I noticed that product id values have strong relation to category_id. \nHere's the histogram of top 10 frequent items by product id (X-axis label 0.5 means product id 5,000,000, 2.0 is 20,000,000, etc)\n\n![enter image description here][1]\n\nFrom the histogram, you can find Category_id 1000010653 is very frequent in low product _id values and not much in 100k to 1,000k range. \n\nSomeone will figure out this eventually(or some are using it already) and use this information, so I guess it is better to disclose early and get some feedback from Kaggle admin. \n\nUsing this information is not beneficial for both competitors and competition host since it does not give any educational value and commercial value. So I hope that the administrator bans use of this information - at least for prize winner (it is practically impossible to ban for others, so it is up to each participants conscience anyway.  :) )\n\nBTW, if you are splitting the train set into several records by product id sequentially, you'd better shuffle it first.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/233979/7722/info_leak.png",
    "233980": "Code for above histogram is:\n\n    # Assume train_df is dataframe of _id and category_id\n    freq_df = pd.crosstab(index=train_df['category_id'], columns='count').sort_values('count', axis=0, ascending=False)\n    top10s = dict([(i, train_df[df['category_id'] == i].index) for i in freq_df.index[:10]])\n    df = pd.DataFrame.from_dict(top10s, orient='index').T\n    df.plot.hist(stacked=True, bins=100, figsize=(16, 6))",
    "234025": "",
    "234027": "You are noticing some serious problem, that could be utilized as cheating. Luckily we do split it randomly. I am doubting if the undisclosed data are as same pattern as your histogram. Did you test that?",
    "234031": "You can see some missing _id in train set, and they are in test set. I believe that test set's _id also follows same pattern as train set's _id.",
    "234038": "&gt;&gt; missing _id in train set, and they are in test set\n\ndoes it also means that if train_product_id is close to test_product_id, their label would be close? It may not be only for Category_id 1000010653?",
    "234039": "From your chart, actually, only 1/10, say,  `product_id 10000010653` have the problem you described, others are a uniform distribution. So it that really a serious problem?",
    "234040": "I think the answer is not. Actually, in the training set, close `product_id` doesn't have close `category_id`. Why do you think so?",
    "234063": "I can say brown category doesn't occur at _id between 2e6 and 6e6 and purple one doesn't occur between 3e6 and 6e6, and this pattern applies to 8 of 10 frequent items. I applied this info to my local validation result, but I see improvements on only really tiny samples(0.0001%).  \n\nFortunately, I believe that this will not be much an issue for this competition :) \n\nBut it would be good to hear some feedback from the host or admin.",
    "234064": "Heng \nYes. Probability of having same label is higher if train/test product _id value is close."
  },
  "source": "meta"
}