{
  "id": 56309,
  "title": "Did anyone get a better result with pseudo-labeling? ",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56309",
  "author_name": "",
  "post_date": "2018-05-08T11:52:37.000976600Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I tried to use soft-labeling on full test data but got a much worse result.  Just wonder if anyone succeeded with pseudo-labeling. The rank doesn't matter as long as the pseudo-labeling helps. A related discussion is here (<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52782#320649\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52782#320649</a>). Thanks.</p>",
  "messages": [
    {
      "id": "325415",
      "postDate": "05/08/2018 11:52:37",
      "content": "<p>I tried to use soft-labeling on full test data but got a much worse result.  Just wonder if anyone succeeded with pseudo-labeling. The rank doesn't matter as long as the pseudo-labeling helps. A related discussion is here (<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52782#320649\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52782#320649</a>). Thanks.</p>",
      "rawMarkdown": "I tried to use soft-labeling on full test data but got a much worse result.  Just wonder if anyone succeeded with pseudo-labeling. The rank doesn't matter as long as the pseudo-labeling helps. A related discussion is here ([https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52782#320649][1]). Thanks.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52782#320649",
      "votes": null
    },
    {
      "id": "325417",
      "postDate": "05/08/2018 11:54:33",
      "content": "<p>I tried, it wasn't working at all for me.</p>",
      "rawMarkdown": "I tried, it wasn't working at all for me.",
      "votes": null
    },
    {
      "id": "325419",
      "postDate": "05/08/2018 11:55:43",
      "content": "<p>I did, from 0.9806 to 0.9816.\nI wrote in detail here :\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56304\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56304</a></p>",
      "rawMarkdown": "I did, from 0.9806 to 0.9816.\nI wrote in detail here :\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56304",
      "votes": null
    },
    {
      "id": "325428",
      "postDate": "05/08/2018 12:02:16",
      "content": "<p>@Amirh, this is not pseudo labeling.  What you did is interesting indeed, I will try it next time, but pseudo labeling is to use make predictions on the test data (which you did), then turn these predictions on labels for test data, then use that labelled data as additional training data.  I don't think you did the latter part from your write up.</p>",
      "rawMarkdown": "Amirh, this is not pseudo labeling.  What you did is interesting indeed, I will try it next time, but pseudo labeling is to use make predictions on the test data (which you did), then turn these predictions on labels for test data, then use that labelled data as additional training data.  I don't think you did the latter part from your write up.",
      "votes": null
    },
    {
      "id": "325431",
      "postDate": "05/08/2018 12:05:22",
      "content": "<p>@ArimH I really like the target encoding idea. </p>",
      "rawMarkdown": "ArimH I really like the target encoding idea.",
      "votes": null
    },
    {
      "id": "325433",
      "postDate": "05/08/2018 12:06:56",
      "content": "<p>@CPMP. I did exactly what you described on pseudo labeling.</p>",
      "rawMarkdown": "CPMP. I did exactly what you described on pseudo labeling.",
      "votes": null
    },
    {
      "id": "325439",
      "postDate": "05/08/2018 12:18:36",
      "content": "<p>tried pseudo labelling and also pseudo labelled target encoding, in which none worked for me, for the latter particularly should be because I didnt follow a leakless target encoding strategy like what AmirH did. I tried generating target encoding from all training data, train with data before day9 4am validate on day 9 4am to 3pm.  Next step is generate target encoding from full training data + test data with pseudo labelling. Train on full training set with best iteration from previous training iterations, predict on test set. It got lesser by 0.004+ in lb</p>",
      "rawMarkdown": "tried pseudo labelling and also pseudo labelled target encoding, in which none worked for me, for the latter particularly should be because I didnt follow a leakless target encoding strategy like what AmirH did. I tried generating target encoding from all training data, train with data before day9 4am validate on day 9 4am to 3pm.  Next step is generate target encoding from full training data + test data with pseudo labelling. Train on full training set with best iteration from previous training iterations, predict on test set. It got lesser by 0.004+ in lb",
      "votes": null
    },
    {
      "id": "325461",
      "postDate": "05/08/2018 12:44:02",
      "content": "<p>@Ee Kin,  thanks for the note.</p>",
      "rawMarkdown": "Ee Kin,  thanks for the note.",
      "votes": null
    },
    {
      "id": "325488",
      "postDate": "05/08/2018 13:17:36",
      "content": "<p>@CPMP \nMy mistake, thanks for the clarification! </p>",
      "rawMarkdown": "CPMP \nMy mistake, thanks for the clarification!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 325417,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/08/2018 11:54:33",
      "content": "<p>I tried, it wasn't working at all for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 325419,
      "author_name": "tpthegreat",
      "author_url": "",
      "post_date": "05/08/2018 11:55:43",
      "content": "<p>I did, from 0.9806 to 0.9816.\nI wrote in detail here :\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56304\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56304</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 325428,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/08/2018 12:02:16",
          "content": "<p>@Amirh, this is not pseudo labeling.  What you did is interesting indeed, I will try it next time, but pseudo labeling is to use make predictions on the test data (which you did), then turn these predictions on labels for test data, then use that labelled data as additional training data.  I don't think you did the latter part from your write up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325431,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "05/08/2018 12:05:22",
          "content": "<p>@ArimH I really like the target encoding idea. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325433,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "05/08/2018 12:06:56",
          "content": "<p>@CPMP. I did exactly what you described on pseudo labeling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 325488,
          "author_name": "tpthegreat",
          "author_url": "",
          "post_date": "05/08/2018 13:17:36",
          "content": "<p>@CPMP \nMy mistake, thanks for the clarification! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 325439,
      "author_name": "dicksonchin93",
      "author_url": "",
      "post_date": "05/08/2018 12:18:36",
      "content": "<p>tried pseudo labelling and also pseudo labelled target encoding, in which none worked for me, for the latter particularly should be because I didnt follow a leakless target encoding strategy like what AmirH did. I tried generating target encoding from all training data, train with data before day9 4am validate on day 9 4am to 3pm.  Next step is generate target encoding from full training data + test data with pseudo labelling. Train on full training set with best iteration from previous training iterations, predict on test set. It got lesser by 0.004+ in lb</p>",
      "votes": null,
      "replies": [
        {
          "id": 325461,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "05/08/2018 12:44:02",
          "content": "<p>@Ee Kin,  thanks for the note.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "325415": "I tried to use soft-labeling on full test data but got a much worse result.  Just wonder if anyone succeeded with pseudo-labeling. The rank doesn't matter as long as the pseudo-labeling helps. A related discussion is here ([https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52782#320649][1]). Thanks.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52782#320649",
    "325417": "I tried, it wasn't working at all for me.",
    "325419": "I did, from 0.9806 to 0.9816.\nI wrote in detail here :\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56304",
    "325428": "Amirh, this is not pseudo labeling.  What you did is interesting indeed, I will try it next time, but pseudo labeling is to use make predictions on the test data (which you did), then turn these predictions on labels for test data, then use that labelled data as additional training data.  I don't think you did the latter part from your write up.",
    "325431": "ArimH I really like the target encoding idea.",
    "325433": "CPMP. I did exactly what you described on pseudo labeling.",
    "325439": "tried pseudo labelling and also pseudo labelled target encoding, in which none worked for me, for the latter particularly should be because I didnt follow a leakless target encoding strategy like what AmirH did. I tried generating target encoding from all training data, train with data before day9 4am validate on day 9 4am to 3pm.  Next step is generate target encoding from full training data + test data with pseudo labelling. Train on full training set with best iteration from previous training iterations, predict on test set. It got lesser by 0.004+ in lb",
    "325461": "Ee Kin,  thanks for the note.",
    "325488": "CPMP \nMy mistake, thanks for the clarification!"
  },
  "source": "meta"
}