{
  "id": 55677,
  "title": "Maybe the shakeup is caused from the duplicated data..",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55677",
  "author_name": "",
  "post_date": "2018-04-30T14:52:01.769363800Z",
  "votes": 36,
  "comment_count": 9,
  "views": 0,
  "content": "<p>We changed the prediction to 0 except the last one if the data is duplicated ，the public LB from 0.9827 to 0.9828......</p>",
  "messages": [
    {
      "id": "321086",
      "postDate": "04/30/2018 14:52:01",
      "content": "<p>We changed the prediction to 0 except the last one if the data is duplicated ，the public LB from 0.9827 to 0.9828......</p>",
      "rawMarkdown": "We changed the prediction to 0 except the last one if the data is duplicated ，the public LB from 0.9827 to 0.9828......",
      "votes": null
    },
    {
      "id": "321092",
      "postDate": "04/30/2018 15:06:05",
      "content": "<p>I'd like to add that we don't know whether this post-processing will be useful for the private leaderboard. Actually, we cannot extract any meaningful information from the duplicated rows.</p>",
      "rawMarkdown": "I'd like to add that we don't know whether this post-processing will be useful for the private leaderboard. Actually, we cannot extract any meaningful information from the duplicated rows.",
      "votes": null
    },
    {
      "id": "321107",
      "postDate": "04/30/2018 16:10:43",
      "content": "<p>I already noticed this 'duplicate' problem 10 days ago and tried hard to know the correct order of duplicates in test data, but can't find the correct one. I'm 100% sure that if someone detects the correct order, he will definitely win and the competition becomes meaningless. I think kaggle admin should share the information (e.g. index of test data in test_supplement) as soon as possible.</p>",
      "rawMarkdown": "I already noticed this 'duplicate' problem 10 days ago and tried hard to know the correct order of duplicates in test data, but can't find the correct one. I'm 100% sure that if someone detects the correct order, he will definitely win and the competition becomes meaningless. I think kaggle admin should share the information (e.g. index of test data in test_supplement) as soon as possible.",
      "votes": null
    },
    {
      "id": "321113",
      "postDate": "04/30/2018 16:26:31",
      "content": "<p>Yes,I agree with you,I think the duplicates is meaningless.Or the organizers should drop duplicates,and the label is the max of the original label.</p>",
      "rawMarkdown": "Yes,I agree with you,I think the duplicates is meaningless.Or the organizers should drop duplicates,and the label is the max of the original label.",
      "votes": null
    },
    {
      "id": "321238",
      "postDate": "04/30/2018 21:32:06",
      "content": "<p>I noticed the issue with duplicate having different target values.  Intuitively the last of duplicates should receive the download,  but this is not true in train data.  From what you say it seems to be the case for test data.  I'll try it.</p>",
      "rawMarkdown": "I noticed the issue with duplicate having different target values.  Intuitively the last of duplicates should receive the download,  but this is not true in train data.  From what you say it seems to be the case for test data.  I'll try it.",
      "votes": null
    },
    {
      "id": "321288",
      "postDate": "04/30/2018 23:49:50",
      "content": "<p>yeah, we want to win against your team with meaningfullness... :)</p>",
      "rawMarkdown": "yeah, we want to win against your team with meaningfullness... :)",
      "votes": null
    },
    {
      "id": "321301",
      "postDate": "05/01/2018 01:04:49",
      "content": "<p><a href=\"/mamasinkgs\">@mamasinkgs</a>, would you elaborate on how knowing the correct order of duplicates will ensure a win?  Kind of new to this, so maybe I'm missing something.</p>",
      "rawMarkdown": "mamasinkgs, would you elaborate on how knowing the correct order of duplicates will ensure a win?  Kind of new to this, so maybe I'm missing something.",
      "votes": null
    },
    {
      "id": "321377",
      "postDate": "05/01/2018 06:03:02",
      "content": "<p>@Matthew, See this discussion <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54229#312231\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54229#312231</a></p>\n\n<p>There are many duplicated rows in train if you ignore <code>is_attributed</code> and <code>attributed_time</code>.  For these rows, one of the duplicate has id_attributed at 1 and the other at 0.  </p>\n\n<p>In test there are also duplicated rows.  If we can predict which of the duplicate is a 0 then we will improve our score.</p>",
      "rawMarkdown": "Matthew, See this discussion https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54229#312231\n\nThere are many duplicated rows in train if you ignore `is_attributed` and `attributed_time`.  For these rows, one of the duplicate has id_attributed at 1 and the other at 0.  \n\nIn test there are also duplicated rows.  If we can predict which of the duplicate is a 0 then we will improve our score.",
      "votes": null
    },
    {
      "id": "321568",
      "postDate": "05/01/2018 15:22:33",
      "content": "<p>I guess the problem with duplicates comes down to : have the samples been sorted using milliseconds or seconds. The former would probably be in favor of having the last duplicate as a one (although all may be zeros). The latter may just be random.</p>\n\n<p>The next question is why time is in seconds only ;-) </p>",
      "rawMarkdown": "I guess the problem with duplicates comes down to : have the samples been sorted using milliseconds or seconds. The former would probably be in favor of having the last duplicate as a one (although all may be zeros). The latter may just be random.\n\nThe next question is why time is in seconds only ;-)",
      "votes": null
    },
    {
      "id": "323177",
      "postDate": "05/04/2018 14:54:07",
      "content": "<p>now That's too late, we just hope it makes no differences in private LB...</p>",
      "rawMarkdown": "now That's too late, we just hope it makes no differences in private LB...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 321092,
      "author_name": "panfeiyang",
      "author_url": "",
      "post_date": "04/30/2018 15:06:05",
      "content": "<p>I'd like to add that we don't know whether this post-processing will be useful for the private leaderboard. Actually, we cannot extract any meaningful information from the duplicated rows.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 321107,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "04/30/2018 16:10:43",
      "content": "<p>I already noticed this 'duplicate' problem 10 days ago and tried hard to know the correct order of duplicates in test data, but can't find the correct one. I'm 100% sure that if someone detects the correct order, he will definitely win and the competition becomes meaningless. I think kaggle admin should share the information (e.g. index of test data in test_supplement) as soon as possible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 321113,
          "author_name": "plantsgo",
          "author_url": "",
          "post_date": "04/30/2018 16:26:31",
          "content": "<p>Yes,I agree with you,I think the duplicates is meaningless.Or the organizers should drop duplicates,and the label is the max of the original label.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321288,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "04/30/2018 23:49:50",
          "content": "<p>yeah, we want to win against your team with meaningfullness... :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321301,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "05/01/2018 01:04:49",
          "content": "<p><a href=\"/mamasinkgs\">@mamasinkgs</a>, would you elaborate on how knowing the correct order of duplicates will ensure a win?  Kind of new to this, so maybe I'm missing something.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321377,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/01/2018 06:03:02",
          "content": "<p>@Matthew, See this discussion <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54229#312231\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54229#312231</a></p>\n\n<p>There are many duplicated rows in train if you ignore <code>is_attributed</code> and <code>attributed_time</code>.  For these rows, one of the duplicate has id_attributed at 1 and the other at 0.  </p>\n\n<p>In test there are also duplicated rows.  If we can predict which of the duplicate is a 0 then we will improve our score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321238,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/30/2018 21:32:06",
      "content": "<p>I noticed the issue with duplicate having different target values.  Intuitively the last of duplicates should receive the download,  but this is not true in train data.  From what you say it seems to be the case for test data.  I'll try it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 321568,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "05/01/2018 15:22:33",
      "content": "<p>I guess the problem with duplicates comes down to : have the samples been sorted using milliseconds or seconds. The former would probably be in favor of having the last duplicate as a one (although all may be zeros). The latter may just be random.</p>\n\n<p>The next question is why time is in seconds only ;-) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 323177,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "05/04/2018 14:54:07",
      "content": "<p>now That's too late, we just hope it makes no differences in private LB...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "321086": "We changed the prediction to 0 except the last one if the data is duplicated ，the public LB from 0.9827 to 0.9828......",
    "321092": "I'd like to add that we don't know whether this post-processing will be useful for the private leaderboard. Actually, we cannot extract any meaningful information from the duplicated rows.",
    "321107": "I already noticed this 'duplicate' problem 10 days ago and tried hard to know the correct order of duplicates in test data, but can't find the correct one. I'm 100% sure that if someone detects the correct order, he will definitely win and the competition becomes meaningless. I think kaggle admin should share the information (e.g. index of test data in test_supplement) as soon as possible.",
    "321113": "Yes,I agree with you,I think the duplicates is meaningless.Or the organizers should drop duplicates,and the label is the max of the original label.",
    "321238": "I noticed the issue with duplicate having different target values.  Intuitively the last of duplicates should receive the download,  but this is not true in train data.  From what you say it seems to be the case for test data.  I'll try it.",
    "321288": "yeah, we want to win against your team with meaningfullness... :)",
    "321301": "mamasinkgs, would you elaborate on how knowing the correct order of duplicates will ensure a win?  Kind of new to this, so maybe I'm missing something.",
    "321377": "Matthew, See this discussion https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54229#312231\n\nThere are many duplicated rows in train if you ignore `is_attributed` and `attributed_time`.  For these rows, one of the duplicate has id_attributed at 1 and the other at 0.  \n\nIn test there are also duplicated rows.  If we can predict which of the duplicate is a 0 then we will improve our score.",
    "321568": "I guess the problem with duplicates comes down to : have the samples been sorted using milliseconds or seconds. The former would probably be in favor of having the last duplicate as a one (although all may be zeros). The latter may just be random.\n\nThe next question is why time is in seconds only ;-)",
    "323177": "now That's too late, we just hope it makes no differences in private LB..."
  },
  "source": "meta"
}