{
  "id": 52782,
  "title": "Is this a good competition for pseudo-labeling?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52782",
  "author_name": "",
  "post_date": "2018-03-23T03:45:03.926608300Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I'm trying really hard to not get sucked into another Kaggle competition, because Kaggle is addictive, but I've already done three competitions in a row and I need to take a break. So one way to motivate myself not to compete here is to just share my idea so I can't benefit from it. (Sorry but not sorry if other people were hoarding this idea to themselves and I blew your cover. Also sorry but not sorry if this idea is bad and makes your models even more overfit and you lose.)</p>\n\n<p>Basically, I see a good deal here that looks like the Jigsaw toxic Kaggle competition that just finished -- really high AUCs and questionable train / test set differences. I was really surprised there to to see the <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\">first place winner</a> winner do so well by using their best ensemble to label the test set data, predict those labels, and then repeat until convergence. Thinking about it, it definitely makes sense as a strategy, especially when your AUCs are high. Given that we also have an extra large supplemental test set, this could be really useful.</p>",
  "messages": [
    {
      "id": "301656",
      "postDate": "03/23/2018 03:45:03",
      "content": "<p>I'm trying really hard to not get sucked into another Kaggle competition, because Kaggle is addictive, but I've already done three competitions in a row and I need to take a break. So one way to motivate myself not to compete here is to just share my idea so I can't benefit from it. (Sorry but not sorry if other people were hoarding this idea to themselves and I blew your cover. Also sorry but not sorry if this idea is bad and makes your models even more overfit and you lose.)</p>\n\n<p>Basically, I see a good deal here that looks like the Jigsaw toxic Kaggle competition that just finished -- really high AUCs and questionable train / test set differences. I was really surprised there to to see the <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\">first place winner</a> winner do so well by using their best ensemble to label the test set data, predict those labels, and then repeat until convergence. Thinking about it, it definitely makes sense as a strategy, especially when your AUCs are high. Given that we also have an extra large supplemental test set, this could be really useful.</p>",
      "rawMarkdown": "I'm trying really hard to not get sucked into another Kaggle competition, because Kaggle is addictive, but I've already done three competitions in a row and I need to take a break. So one way to motivate myself not to compete here is to just share my idea so I can't benefit from it. (Sorry but not sorry if other people were hoarding this idea to themselves and I blew your cover. Also sorry but not sorry if this idea is bad and makes your models even more overfit and you lose.)\n\nBasically, I see a good deal here that looks like the Jigsaw toxic Kaggle competition that just finished -- really high AUCs and questionable train / test set differences. I was really surprised there to to see the [first place winner](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557) winner do so well by using their best ensemble to label the test set data, predict those labels, and then repeat until convergence. Thinking about it, it definitely makes sense as a strategy, especially when your AUCs are high. Given that we also have an extra large supplemental test set, this could be really useful.",
      "votes": null
    },
    {
      "id": "301689",
      "postDate": "03/23/2018 05:41:24",
      "content": "<p>Put in to do list. Thanks!</p>\n\n<p>So according to them this is make sense when there is a difference in train and test distributions?</p>",
      "rawMarkdown": "Put in to do list. Thanks!\n\nSo according to them this is make sense when there is a difference in train and test distributions?",
      "votes": null
    },
    {
      "id": "302105",
      "postDate": "03/23/2018 17:39:31",
      "content": "<p>It seems theoretically to make sense whenever you can guess test labels with really high accuracy.</p>",
      "rawMarkdown": "It seems theoretically to make sense whenever you can guess test labels with really high accuracy.",
      "votes": null
    },
    {
      "id": "302179",
      "postDate": "03/23/2018 19:12:50",
      "content": "<p>Peter, even if it's early to say I wouldn't bet on it; </p>\n\n<p><strong>Pseudo-labeling is a trade off</strong> , its effect  on one hand depends on accuracy , as you have said. But just as important, it depends on how much you gain when you add data to train.</p>\n\n<p>In toxic competition there was not so much data for the complexity of the task, many symptoms apart of the number of less frequent labels, pretrained embeddings worked better than custom, high number of out of voc... Good setup for pseudo -labeling. </p>\n\n<p>I entered toxic really late so a ton of things I didn't try but pseudolabelling was definitely high in  my list then. In this one its not even in the list. But I could be wrong...</p>\n\n<p>Anyway... c'mon man, it's not too late,  why not join the party? Its not such a bad adiction! :)</p>",
      "rawMarkdown": "Peter, even if it's early to say I wouldn't bet on it; \n\n**Pseudo-labeling is a trade off** , its effect  on one hand depends on accuracy , as you have said. But just as important, it depends on how much you gain when you add data to train.\n\nIn toxic competition there was not so much data for the complexity of the task, many symptoms apart of the number of less frequent labels, pretrained embeddings worked better than custom, high number of out of voc... Good setup for pseudo -labeling. \n\nI entered toxic really late so a ton of things I didn't try but pseudolabelling was definitely high in  my list then. In this one its not even in the list. But I could be wrong...\n\nAnyway... c'mon man, it's not too late,  why not join the party? Its not such a bad adiction! :)",
      "votes": null
    },
    {
      "id": "302190",
      "postDate": "03/23/2018 19:27:15",
      "content": "<blockquote>\n  <p>it's not too late, why not join the party? Its not such a bad adiction! :)</p>\n</blockquote>\n\n<p>I tell myself that a lot... there are much worse things to have an addiction to. :) But I've been making doing well on Kaggle pretty much my #1 priority for the past five months now. I'm very delighted with the results I've shown for it on Kaggle, but I need to pick up slack on the things I've been neglecting.</p>\n\n<p>But don't worry, I'll rejoin y'all soon enough! It's going to be a long road to Grandmaster. :)</p>",
      "rawMarkdown": "&gt; it's not too late, why not join the party? Its not such a bad adiction! :)\n\nI tell myself that a lot... there are much worse things to have an addiction to. :) But I've been making doing well on Kaggle pretty much my #1 priority for the past five months now. I'm very delighted with the results I've shown for it on Kaggle, but I need to pick up slack on the things I've been neglecting.\n\nBut don't worry, I'll rejoin y'all soon enough! It's going to be a long road to Grandmaster. :)",
      "votes": null
    },
    {
      "id": "302194",
      "postDate": "03/23/2018 19:28:58",
      "content": "<p>&gt; In toxic competition there was not so much data for the complexity of the task, many symptoms apart of the number of less frequent labels, pretrained embeddings worked better than custom, high number of out of voc... Good setup for pseudo -labeling.</p>\n\n<p>That's really interesting. I didn't get a sense here that the data size was that different from toxic, but the fact that this competition is not text-based is definitely a huge difference.</p>",
      "rawMarkdown": "&gt; In toxic competition there was not so much data for the complexity of the task, many symptoms apart of the number of less frequent labels, pretrained embeddings worked better than custom, high number of out of voc... Good setup for pseudo -labeling.\n\nThat's really interesting. I didn't get a sense here that the data size was that different from toxic, but the fact that this competition is not text-based is definitely a huge difference.",
      "votes": null
    },
    {
      "id": "320594",
      "postDate": "04/29/2018 07:10:43",
      "content": "<p>I agree with miguel on this. The number of samples on here is massive in comparison to the toxic comment challenge. Seems like that extra bit of single you might be able to suss out would either be very small since you already have millions of samples to already learn from or simply incorrect because the current results aren't really all that good. </p>",
      "rawMarkdown": "I agree with miguel on this. The number of samples on here is massive in comparison to the toxic comment challenge. Seems like that extra bit of single you might be able to suss out would either be very small since you already have millions of samples to already learn from or simply incorrect because the current results aren't really all that good.",
      "votes": null
    },
    {
      "id": "320649",
      "postDate": "04/29/2018 11:55:40",
      "content": "<p><a href=\"/peterhurford\">@peterhurford</a> only a few days to go, so the damage is limited, dive into it and have some fun :p</p>",
      "rawMarkdown": "peterhurford only a few days to go, so the damage is limited, dive into it and have some fun :p",
      "votes": null
    },
    {
      "id": "320723",
      "postDate": "04/29/2018 16:11:39",
      "content": "<p>Thanks! I'm going to do the <a href=\"https://www.kaggle.com/c/avito-demand-prediction/\">Avito Demand Competition</a> though. :)</p>",
      "rawMarkdown": "Thanks! I'm going to do the [Avito Demand Competition](https://www.kaggle.com/c/avito-demand-prediction/) though. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 301689,
      "author_name": "muhammadalfiansyah",
      "author_url": "",
      "post_date": "03/23/2018 05:41:24",
      "content": "<p>Put in to do list. Thanks!</p>\n\n<p>So according to them this is make sense when there is a difference in train and test distributions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 302105,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "03/23/2018 17:39:31",
          "content": "<p>It seems theoretically to make sense whenever you can guess test labels with really high accuracy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 302179,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "03/23/2018 19:12:50",
      "content": "<p>Peter, even if it's early to say I wouldn't bet on it; </p>\n\n<p><strong>Pseudo-labeling is a trade off</strong> , its effect  on one hand depends on accuracy , as you have said. But just as important, it depends on how much you gain when you add data to train.</p>\n\n<p>In toxic competition there was not so much data for the complexity of the task, many symptoms apart of the number of less frequent labels, pretrained embeddings worked better than custom, high number of out of voc... Good setup for pseudo -labeling. </p>\n\n<p>I entered toxic really late so a ton of things I didn't try but pseudolabelling was definitely high in  my list then. In this one its not even in the list. But I could be wrong...</p>\n\n<p>Anyway... c'mon man, it's not too late,  why not join the party? Its not such a bad adiction! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 302190,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "03/23/2018 19:27:15",
          "content": "<blockquote>\n  <p>it's not too late, why not join the party? Its not such a bad adiction! :)</p>\n</blockquote>\n\n<p>I tell myself that a lot... there are much worse things to have an addiction to. :) But I've been making doing well on Kaggle pretty much my #1 priority for the past five months now. I'm very delighted with the results I've shown for it on Kaggle, but I need to pick up slack on the things I've been neglecting.</p>\n\n<p>But don't worry, I'll rejoin y'all soon enough! It's going to be a long road to Grandmaster. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 302194,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "03/23/2018 19:28:58",
          "content": "<p>&gt; In toxic competition there was not so much data for the complexity of the task, many symptoms apart of the number of less frequent labels, pretrained embeddings worked better than custom, high number of out of voc... Good setup for pseudo -labeling.</p>\n\n<p>That's really interesting. I didn't get a sense here that the data size was that different from toxic, but the fact that this competition is not text-based is definitely a huge difference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320594,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "04/29/2018 07:10:43",
          "content": "<p>I agree with miguel on this. The number of samples on here is massive in comparison to the toxic comment challenge. Seems like that extra bit of single you might be able to suss out would either be very small since you already have millions of samples to already learn from or simply incorrect because the current results aren't really all that good. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320649,
      "author_name": "yifanxie",
      "author_url": "",
      "post_date": "04/29/2018 11:55:40",
      "content": "<p><a href=\"/peterhurford\">@peterhurford</a> only a few days to go, so the damage is limited, dive into it and have some fun :p</p>",
      "votes": null,
      "replies": [
        {
          "id": 320723,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "04/29/2018 16:11:39",
          "content": "<p>Thanks! I'm going to do the <a href=\"https://www.kaggle.com/c/avito-demand-prediction/\">Avito Demand Competition</a> though. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "301656": "I'm trying really hard to not get sucked into another Kaggle competition, because Kaggle is addictive, but I've already done three competitions in a row and I need to take a break. So one way to motivate myself not to compete here is to just share my idea so I can't benefit from it. (Sorry but not sorry if other people were hoarding this idea to themselves and I blew your cover. Also sorry but not sorry if this idea is bad and makes your models even more overfit and you lose.)\n\nBasically, I see a good deal here that looks like the Jigsaw toxic Kaggle competition that just finished -- really high AUCs and questionable train / test set differences. I was really surprised there to to see the [first place winner](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557) winner do so well by using their best ensemble to label the test set data, predict those labels, and then repeat until convergence. Thinking about it, it definitely makes sense as a strategy, especially when your AUCs are high. Given that we also have an extra large supplemental test set, this could be really useful.",
    "301689": "Put in to do list. Thanks!\n\nSo according to them this is make sense when there is a difference in train and test distributions?",
    "302105": "It seems theoretically to make sense whenever you can guess test labels with really high accuracy.",
    "302179": "Peter, even if it's early to say I wouldn't bet on it; \n\n**Pseudo-labeling is a trade off** , its effect  on one hand depends on accuracy , as you have said. But just as important, it depends on how much you gain when you add data to train.\n\nIn toxic competition there was not so much data for the complexity of the task, many symptoms apart of the number of less frequent labels, pretrained embeddings worked better than custom, high number of out of voc... Good setup for pseudo -labeling. \n\nI entered toxic really late so a ton of things I didn't try but pseudolabelling was definitely high in  my list then. In this one its not even in the list. But I could be wrong...\n\nAnyway... c'mon man, it's not too late,  why not join the party? Its not such a bad adiction! :)",
    "302190": "&gt; it's not too late, why not join the party? Its not such a bad adiction! :)\n\nI tell myself that a lot... there are much worse things to have an addiction to. :) But I've been making doing well on Kaggle pretty much my #1 priority for the past five months now. I'm very delighted with the results I've shown for it on Kaggle, but I need to pick up slack on the things I've been neglecting.\n\nBut don't worry, I'll rejoin y'all soon enough! It's going to be a long road to Grandmaster. :)",
    "302194": "&gt; In toxic competition there was not so much data for the complexity of the task, many symptoms apart of the number of less frequent labels, pretrained embeddings worked better than custom, high number of out of voc... Good setup for pseudo -labeling.\n\nThat's really interesting. I didn't get a sense here that the data size was that different from toxic, but the fact that this competition is not text-based is definitely a huge difference.",
    "320594": "I agree with miguel on this. The number of samples on here is massive in comparison to the toxic comment challenge. Seems like that extra bit of single you might be able to suss out would either be very small since you already have millions of samples to already learn from or simply incorrect because the current results aren't really all that good.",
    "320649": "peterhurford only a few days to go, so the damage is limited, dive into it and have some fun :p",
    "320723": "Thanks! I'm going to do the [Avito Demand Competition](https://www.kaggle.com/c/avito-demand-prediction/) though. :)"
  },
  "source": "meta"
}