{
  "id": 160333,
  "title": "Jigsaw Competition attracts 1300+ less teams every year",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/160333",
  "author_name": "YaGana Sheriff-Hussaini",
  "post_date": "2020-06-20T21:04:28.105000",
  "votes": 8,
  "comment_count": 43,
  "views": 0,
  "content": "<p>I noticed that the number of teams in this comeptition is a lot lower than last year's which prompted me to check all the numbers. I thought it is interesting that the number of team decrease is roughly the same yearly plus some decay factor😃 . See below:-</p>\n\n<ol>\n<li><p><a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge\">Toxic Comment Classification Challenge</a> - total participation is 4,550 teams.</p></li>\n<li><p><a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification\">Jigsaw Unintended Bias in Toxicity Classification</a> - total participation is 3,165 teams (down by 1385)</p></li>\n<li><p><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/leaderboard\">Jigsaw Multilingual Toxic Comment Classification</a> - total participation is 1,480 teams (down by 1685). This one may go down on closing after cheaters removal.</p></li>\n</ol>\n\n<h3>UPDATE:-</h3>\n\n<p>I am updating my main post with a summary of the most voted comments for anyone who did not want to go through them.</p>\n\n<ol>\n<li>Another very engaging and arguably even more interesting NLP competition going on at the same time i.e. <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction\">Tweet Sentiment Extraction</a>.</li>\n<li>In the past we had unlimited GPU time for everyone, no external data except pre-provided vectors, people developing elegant simple models, and in the end nice solutions. <a href=\"/philippsinger\">@philippsinger</a> </li>\n<li><ul><li>jigsaw1: we can use CPU to win. - jigsaw2: we can use free kaggle GPU(no time limit). - jigsaw3: we only have 30 hours per week <a href=\"/mcggood\">@mcggood</a> </li></ul></li>\n<li>Needing to use TPU's to even have the ability to train these massive models is a large barrier. It was easy when it was TF-IDF and LSTM. Complexity and difficulty seems to keep going up. It is not very easy for a newcomer to jump in from zero on most of these competitions. <a href=\"/ryches\">@ryches</a> </li>\n</ol>",
  "messages": [
    {
      "id": 894877,
      "postDate": "2020-06-20T21:04:28.107Z",
      "content": "<p>I noticed that the number of teams in this comeptition is a lot lower than last year's which prompted me to check all the numbers. I thought it is interesting that the number of team decrease is roughly the same yearly plus some decay factor😃 . See below:-</p>\n\n<ol>\n<li><p><a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge\">Toxic Comment Classification Challenge</a> - total participation is 4,550 teams.</p></li>\n<li><p><a href=\"https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification\">Jigsaw Unintended Bias in Toxicity Classification</a> - total participation is 3,165 teams (down by 1385)</p></li>\n<li><p><a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/leaderboard\">Jigsaw Multilingual Toxic Comment Classification</a> - total participation is 1,480 teams (down by 1685). This one may go down on closing after cheaters removal.</p></li>\n</ol>\n\n<h3>UPDATE:-</h3>\n\n<p>I am updating my main post with a summary of the most voted comments for anyone who did not want to go through them.</p>\n\n<ol>\n<li>Another very engaging and arguably even more interesting NLP competition going on at the same time i.e. <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction\">Tweet Sentiment Extraction</a>.</li>\n<li>In the past we had unlimited GPU time for everyone, no external data except pre-provided vectors, people developing elegant simple models, and in the end nice solutions. <a href=\"/philippsinger\">@philippsinger</a> </li>\n<li><ul><li>jigsaw1: we can use CPU to win. - jigsaw2: we can use free kaggle GPU(no time limit). - jigsaw3: we only have 30 hours per week <a href=\"/mcggood\">@mcggood</a> </li></ul></li>\n<li>Needing to use TPU's to even have the ability to train these massive models is a large barrier. It was easy when it was TF-IDF and LSTM. Complexity and difficulty seems to keep going up. It is not very easy for a newcomer to jump in from zero on most of these competitions. <a href=\"/ryches\">@ryches</a> </li>\n</ol>",
      "rawMarkdown": "I noticed that the number of teams in this comeptition is a lot lower than last year's which prompted me to check all the numbers. I thought it is interesting that the number of team decrease is roughly the same yearly plus some decay factor😃 . See below:-\n\n1. [Toxic Comment Classification Challenge](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge) - total participation is 4,550 teams.\n\n2. [Jigsaw Unintended Bias in Toxicity Classification](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification) - total participation is 3,165 teams (down by 1385)\n\n3. [Jigsaw Multilingual Toxic Comment Classification](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/leaderboard) - total participation is 1,480 teams (down by 1685). This one may go down on closing after cheaters removal.\n\n### UPDATE:-\n\nI am updating my main post with a summary of the most voted comments for anyone who did not want to go through them.\n\n1. Another very engaging and arguably even more interesting NLP competition going on at the same time i.e. [Tweet Sentiment Extraction](https://www.kaggle.com/c/tweet-sentiment-extraction).\n2. In the past we had unlimited GPU time for everyone, no external data except pre-provided vectors, people developing elegant simple models, and in the end nice solutions. @philippsinger \n3. - jigsaw1: we can use CPU to win. - jigsaw2: we can use free kaggle GPU(no time limit). - jigsaw3: we only have 30 hours per week @mcggood \n4. Needing to use TPU's to even have the ability to train these massive models is a large barrier. It was easy when it was TF-IDF and LSTM. Complexity and difficulty seems to keep going up. It is not very easy for a newcomer to jump in from zero on most of these competitions. @ryches \n",
      "votes": 8
    },
    {
      "id": 895882,
      "postDate": "2020-06-21T17:05:17.233Z",
      "content": "<p>It's getting way less accessible. Needing to use TPU's to even have the ability to train these massive models is a large barrier. It was easy when it was TF-IDF and LSTM. I feel like the compute requirement, complexity and difficulty seems to keep going up. It is not very easy for a newcomer to jump in from zero on most of these competitions. </p>",
      "rawMarkdown": "It's getting way less accessible. Needing to use TPU's to even have the ability to train these massive models is a large barrier. It was easy when it was TF-IDF and LSTM. I feel like the compute requirement, complexity and difficulty seems to keep going up. It is not very easy for a newcomer to jump in from zero on most of these competitions. ",
      "votes": 8,
      "replies": [
        {
          "id": 895905,
          "postDate": "2020-06-21T17:19:24.327Z",
          "content": "<p>My first competition was Quora, and to me it was nearly the perfect setup, except of the two stage re-runs which they would not need to do now anylonger having synchronous kernels.</p>\n\n<p>It had limited GPU time for everyone, no external data except pre-provided vectors, people developing elegant simple models, and in the end nice solutions.</p>\n\n<p>Nowadays it is getting more complex to run the models, external data matters a lot often, and everything is getting fuzzy. I personally, really like constrained environments, even though that many people complain about kernel competitions.</p>",
          "rawMarkdown": "My first competition was Quora, and to me it was nearly the perfect setup, except of the two stage re-runs which they would not need to do now anylonger having synchronous kernels.\n\nIt had limited GPU time for everyone, no external data except pre-provided vectors, people developing elegant simple models, and in the end nice solutions.\n\nNowadays it is getting more complex to run the models, external data matters a lot often, and everything is getting fuzzy. I personally, really like constrained environments, even though that many people complain about kernel competitions.",
          "votes": 15
        },
        {
          "id": 895926,
          "postDate": "2020-06-21T17:34:45.777Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>, Quora was my 2nd competition and I was new to NLP as well found that Quora contest  as a fartile ground for learning and experimentation.</p>\n\n<p>UPDATE:</p>\n\n<p>I did not realize you meant the 2nd Quora competition, my post above is refering to the <a href=\"https://www.kaggle.com/c/quora-question-pairs\">Quora Question Pairs</a> competition 😄 </p>",
          "rawMarkdown": "@philippsinger, Quora was my 2nd competition and I was new to NLP as well found that Quora contest  as a fartile ground for learning and experimentation.\n\nUPDATE:\n\nI did not realize you meant the 2nd Quora competition, my post above is refering to the [Quora Question Pairs](https://www.kaggle.com/c/quora-question-pairs) competition 😄 ",
          "votes": 1
        },
        {
          "id": 895931,
          "postDate": "2020-06-21T17:37:34.857Z",
          "content": "<p>I agree that having a constrained environment is nice, but sometimes the creativity that removing those restrictions provides is nice too. In that Quora competition if we were allowed to use whatever we wanted then BERT would have dominated there which makes you wonder if that was really the best solution for the host. It is a bit of an arms race unless someone puts a cap on our solutions. </p>",
          "rawMarkdown": "I agree that having a constrained environment is nice, but sometimes the creativity that removing those restrictions provides is nice too. In that Quora competition if we were allowed to use whatever we wanted then BERT would have dominated there which makes you wonder if that was really the best solution for the host. It is a bit of an arms race unless someone puts a cap on our solutions. ",
          "votes": 1
        },
        {
          "id": 895936,
          "postDate": "2020-06-21T17:41:27.667Z",
          "content": "<p>I really like constrained environments too (This may seem biased given my two gold medals come from Kernel competitions ). Even though Mercari was a bit too constrained IMHO.</p>\n\n<p>But this competition helped me to get familiar with TPUs, and I have to  say I really envoy now training quickly on large text or Images dataset with TPUs. I will be using 100% TPU on SIIM ISIC competition</p>",
          "rawMarkdown": "I really like constrained environments too (This may seem biased given my two gold medals come from Kernel competitions ). Even though Mercari was a bit too constrained IMHO.\n\n \nBut this competition helped me to get familiar with TPUs, and I have to  say I really envoy now training quickly on large text or Images dataset with TPUs. I will be using 100% TPU on SIIM ISIC competition"
        },
        {
          "id": 895950,
          "postDate": "2020-06-21T17:50:47.113Z",
          "content": "<p><a href=\"/ryches\">@ryches</a> The question is what is more valuable? Finetuning Bert or finding simpler models that are on-par with Bert? I think both have its value.</p>",
          "rawMarkdown": "@ryches The question is what is more valuable? Finetuning Bert or finding simpler models that are on-par with Bert? I think both have its value.",
          "votes": 1
        },
        {
          "id": 895994,
          "postDate": "2020-06-21T18:32:44.600Z",
          "content": "<p>It depends on the the value of a more accurate model. I would argue that scraping the last % out of this toxic comment stuff probably isn't really that valuable and a more constrained solution is probably what they want. That is common of a lot of kaggle problems though.</p>\n\n<p>Definitely value to exploring both constrained and unconstrained approaches</p>",
          "rawMarkdown": "It depends on the the value of a more accurate model. I would argue that scraping the last % out of this toxic comment stuff probably isn't really that valuable and a more constrained solution is probably what they want. That is common of a lot of kaggle problems though.\n\nDefinitely value to exploring both constrained and unconstrained approaches",
          "votes": 1
        },
        {
          "id": 896025,
          "postDate": "2020-06-21T18:59:08.193Z",
          "content": "<p>The way I see it is that developing a simple high performing model in a constrained environment challenges me to dig deeper. This is why I prefer that than fine tuning the new kid on the block like BERT.</p>",
          "rawMarkdown": "The way I see it is that developing a simple high performing model in a constrained environment challenges me to dig deeper. This is why I prefer that than fine tuning the new kid on the block like BERT.",
          "votes": 1
        },
        {
          "id": 896060,
          "postDate": "2020-06-21T19:41:12.863Z",
          "content": "<p>However, BERT is easier to parallelize than LSTM...And with the same amount of data , you can  make it run faster than LSTM (by choosing for instance which hidden layers to freeze/fine-tune) while still outperforming LSTM. </p>",
          "rawMarkdown": "However, BERT is easier to parallelize than LSTM...And with the same amount of data , you can  make it run faster than LSTM (by choosing for instance which hidden layers to freeze/fine-tune) while still outperforming LSTM. \n\n"
        },
        {
          "id": 896062,
          "postDate": "2020-06-21T19:43:54.477Z",
          "content": "<p>My issue is that noone ever tries simpler architectures nowadays because you have a bias towards starting with SOTA models. Constrained Kaggle competitions might force you to try other stuff outside of current research, and maybe it would lead to a higher chance of finding new architectures or solutions to the problem at hand.</p>",
          "rawMarkdown": "My issue is that noone ever tries simpler architectures nowadays because you have a bias towards starting with SOTA models. Constrained Kaggle competitions might force you to try other stuff outside of current research, and maybe it would lead to a higher chance of finding new architectures or solutions to the problem at hand.",
          "votes": 6
        },
        {
          "id": 896077,
          "postDate": "2020-06-21T20:02:32.707Z",
          "content": "<p><a href=\"/serigne\">@serigne</a>, agree and I am not against BERT by the way. Various fine tuned and nicely designed RoBERTa models got my team out of trouble from #241 to 41st in the Tweet competition.</p>\n\n<p><a href=\"/philippsinger\">@philippsinger</a> my sentiments exactly. </p>\n\n<p>I just realised the content of my post must be really really boring going by the number of upvotes😃 Having said that, I am so happy it pompted all these opinions I am reading. Thanks guys and girrls.</p>",
          "rawMarkdown": "@serigne, agree and I am not against BERT by the way. Various fine tuned and nicely designed RoBERTa models got my team out of trouble from #241 to 41st in the Tweet competition.\n\n@philippsinger my sentiments exactly. \n\nI just realised the content of my post must be really really boring going by the number of upvotes😃 Having said that, I am so happy it pompted all these opinions I am reading. Thanks guys and girrls.",
          "votes": 1
        },
        {
          "id": 896205,
          "postDate": "2020-06-22T00:53:22.890Z",
          "content": "<p>The problem with artificial constraints is you often end up emphasizing engineering (e.g., working around CPU/memory bottlenecks or runtime limits). That's a useful skillset but not more important than being able to tweak SOTA models. </p>\n\n<p>With academia moving to using TPU pods, model sizes are going to increase. There's also research showing that given a fixed compute budget, it's better to train large models for a short period and compressing them, rather than training small/moderately-sized models for longer. Not a bad idea for a Kaggler to acclimatize themselves to handling large models. </p>\n\n<p>RL and generative problems (e.g., Halite, ARC) are what you're looking for if you want something requiring more creativity. </p>",
          "rawMarkdown": "The problem with artificial constraints is you often end up emphasizing engineering (e.g., working around CPU/memory bottlenecks or runtime limits). That's a useful skillset but not more important than being able to tweak SOTA models. \n\nWith academia moving to using TPU pods, model sizes are going to increase. There's also research showing that given a fixed compute budget, it's better to train large models for a short period and compressing them, rather than training small/moderately-sized models for longer. Not a bad idea for a Kaggler to acclimatize themselves to handling large models. \n\nRL and generative problems (e.g., Halite, ARC) are what you're looking for if you want something requiring more creativity. ",
          "votes": 4
        },
        {
          "id": 896210,
          "postDate": "2020-06-22T01:05:52.297Z",
          "content": "<p><a href=\"/leecming\">@leecming</a> you are quite right about ARC, I was too busy at work to be able to participate at the time but believe that is a glimse into the future of ML. Halite I thought look very interesting from the Kaggle email invitation, I will be taking a look at that after this competition.</p>",
          "rawMarkdown": "@leecming you are quite right about ARC, I was too busy at work to be able to participate at the time but believe that is a glimse into the future of ML. Halite I thought look very interesting from the Kaggle email invitation, I will be taking a look at that after this competition."
        },
        {
          "id": 896658,
          "postDate": "2020-06-22T10:54:59.927Z",
          "content": "<p>Yeah, I totally agree. I'm a newbie to nlp, and this is my first nlp competition. In fact, I just started to learn nlp months ago by cs224n. I spent a lot of time trying to run these super large models on TPUs with pytorch at the beginning. This is painful, but learnt a lot.  Thanks for those excellent public kernels.\nBWT, may I ask what the general parameters you try first while dealing with these large models? Is it like \"lr: 5e-5, 3e-5, 2e-5; epochs: 2, 3, 4\" as suggested by the BERT paper? </p>",
          "rawMarkdown": "Yeah, I totally agree. I'm a newbie to nlp, and this is my first nlp competition. In fact, I just started to learn nlp months ago by cs224n. I spent a lot of time trying to run these super large models on TPUs with pytorch at the beginning. This is painful, but learnt a lot.  Thanks for those excellent public kernels.\nBWT, may I ask what the general parameters you try first while dealing with these large models? Is it like \"lr: 5e-5, 3e-5, 2e-5; epochs: 2, 3, 4\" as suggested by the BERT paper? ",
          "votes": 1
        },
        {
          "id": 900208,
          "postDate": "2020-06-24T17:24:21.393Z",
          "content": "<p><a href=\"/godelscat\">@godelscat</a>, you can try going through some of the starter BERT model kernels. The steps in the following articles is a good one to follow.</p>\n\n<ol>\n<li><a href=\"https://towardsdatascience.com/fine-tuning-bert-for-text-classification-with-farm-2880665065e2\">https://towardsdatascience.com/fine-tuning-bert-for-text-classification-with-farm-2880665065e2</a></li>\n<li><a href=\"https://mccormickml.com/2019/07/22/BERT-fine-tuning/\">https://mccormickml.com/2019/07/22/BERT-fine-tuning/</a></li>\n<li><a href=\"https://medium.com/curation-corporation/fine-tuning-bart-for-abstractive-text-summarisation-with-fastai2-d7a2ad676a13\">https://medium.com/curation-corporation/fine-tuning-bart-for-abstractive-text-summarisation-with-fastai2-d7a2ad676a13</a></li>\n</ol>",
          "rawMarkdown": "@godelscat, you can try going through some of the starter BERT model kernels. The steps in the following articles is a good one to follow.\n\n1. https://towardsdatascience.com/fine-tuning-bert-for-text-classification-with-farm-2880665065e2\n2. https://mccormickml.com/2019/07/22/BERT-fine-tuning/\n3. https://medium.com/curation-corporation/fine-tuning-bart-for-abstractive-text-summarisation-with-fastai2-d7a2ad676a13\n"
        },
        {
          "id": 900450,
          "postDate": "2020-06-24T20:02:59.217Z",
          "content": "<blockquote>\n  <p>My issue is that noone ever tries simpler architectures nowadays</p>\n</blockquote>\n\n<p>I used linear regression in Covid forecasting, and this was competitive with best models once we factor out difference in training data time span.</p>",
          "rawMarkdown": "&gt; My issue is that noone ever tries simpler architectures nowadays\n\nI used linear regression in Covid forecasting, and this was competitive with best models once we factor out difference in training data time span."
        },
        {
          "id": 900601,
          "postDate": "2020-06-24T23:32:30.943Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, yep that is the thing. In fact I trained FM-FTRL, NBSVM, Single-Wordbatch-LGBM, and PooledGRU-Word2Vec-using-Gensim models in this competition that helped my ensemble score. These models scored low in the range of 0.8750 to 0.9000 on the public LB. I did not know what some of them scored until after the competition in late submission because I was submitting the ensemble based on CV.</p>",
          "rawMarkdown": "@cpmpml, yep that is the thing. In fact I trained FM-FTRL, NBSVM, Single-Wordbatch-LGBM, and PooledGRU-Word2Vec-using-Gensim models in this competition that helped my ensemble score. These models scored low in the range of 0.8750 to 0.9000 on the public LB. I did not know what some of them scored until after the competition in late submission because I was submitting the ensemble based on CV."
        },
        {
          "id": 901462,
          "postDate": "2020-06-25T13:53:23.990Z",
          "content": "<blockquote>\n  <p>With academia moving to using TPU pods, , model sizes are going to increase.</p>\n</blockquote>\n\n<p>Is TPU the reason really?</p>\n\n<p>AFAIK, the largest model ever was trained on GPU: <a href=\"https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/\">https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/</a></p>",
          "rawMarkdown": "&gt; With academia moving to using TPU pods, , model sizes are going to increase.\n\nIs TPU the reason really?\n\nAFAIK, the largest model ever was trained on GPU: https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/",
          "votes": 1
        },
        {
          "id": 902075,
          "postDate": "2020-06-25T23:03:17.413Z",
          "content": "<p>GPU too - Nvidia recently announced their A100 GPU with 40GB VRAM. There's an old saying for hard drives about how applications increase in size to fit growing hard drive capacities. Same with VRAM. </p>\n\n<p>I have two RTX Titans at home with a combined 48GB VRAM. Even with gradient accumulation and mixed precision training, I'm having trouble squeezing SOTA NLP models for pretraining tasks e.g. MLM.</p>\n\n<p>I've half thought about buying a DGX for home use. Can you get me an employee discount :P</p>",
          "rawMarkdown": "GPU too - Nvidia recently announced their A100 GPU with 40GB VRAM. There's an old saying for hard drives about how applications increase in size to fit growing hard drive capacities. Same with VRAM. \n\nI have two RTX Titans at home with a combined 48GB VRAM. Even with gradient accumulation and mixed precision training, I'm having trouble squeezing SOTA NLP models for pretraining tasks e.g. MLM.\n\nI've half thought about buying a DGX for home use. Can you get me an employee discount :P",
          "votes": 1
        },
        {
          "id": 902760,
          "postDate": "2020-06-26T10:45:27.900Z",
          "content": "<p>There are no employee discounts for DGX machines.  I guess only few employees could afford these.  And except for DGX Stations, DGX machines are mean to be used in racks in datacenters.  </p>\n\n<p>We can get PCIe GPU cards eg 2080 Ti,  with a small discount, but that's not what you are discussing, right?  And of course we can't resell.</p>\n\n<p>An alternative to V100 and DGX Station that may be worth considering is RTX 8000 cards.  They are PCIe cards, with 48 GB memory. Jason Antic, the deoldify guy just moved to a 4x RTX 8000 machine for instance: <a href=\"https://twitter.com/citnaj/status/1271892395798392832\">https://twitter.com/citnaj/status/1271892395798392832</a></p>\n\n<p>I am puzzled by your comment about SOTA models.  I trained my models on one V100 at a time.  It has 32GB memory, which is less than what you have.  I'm forced to us small batch size, and as you said, gradient accumulation and mixed precision.  But I can train any of the huggingfaces model.</p>\n\n<p>What may be worrisome is if people take advantage of A100 expanded capacity to move SOTA models beyond what V100 can handle.</p>",
          "rawMarkdown": "There are no employee discounts for DGX machines.  I guess only few employees could afford these.  And except for DGX Stations, DGX machines are mean to be used in racks in datacenters.  \n\nWe can get PCIe GPU cards eg 2080 Ti,  with a small discount, but that's not what you are discussing, right?  And of course we can't resell.\n\nAn alternative to V100 and DGX Station that may be worth considering is RTX 8000 cards.  They are PCIe cards, with 48 GB memory. Jason Antic, the deoldify guy just moved to a 4x RTX 8000 machine for instance: https://twitter.com/citnaj/status/1271892395798392832\n\nI am puzzled by your comment about SOTA models.  I trained my models on one V100 at a time.  It has 32GB memory, which is less than what you have.  I'm forced to us small batch size, and as you said, gradient accumulation and mixed precision.  But I can train any of the huggingfaces model.\n\nWhat may be worrisome is if people take advantage of A100 expanded capacity to move SOTA models beyond what V100 can handle.\n",
          "votes": 1
        },
        {
          "id": 902824,
          "postDate": "2020-06-26T11:39:48.637Z",
          "content": "<p>@CPMP but the price of 4 RTX Quadro 8000 begin to be a bit insane right ? 😄 </p>",
          "rawMarkdown": "@CPMP but the price of 4 RTX Quadro 8000 begin to be a bit insane right ? 😄 "
        },
        {
          "id": 903626,
          "postDate": "2020-06-27T02:10:45.883Z",
          "content": "<p>I've actually thought about building a 4x RTX 8000 server but I'm waiting to see retail availability and pricing of the A100 PCIe component. </p>\n\n<p>About not being able to fit the Toxic Comment models in memory, I was referring more to pretraining MLM models. They have a much bigger memory footprint because of their classifier head (shape: seq len, vocab size). I'm going to try to use the electra approach in the future to pretrain - it's supposedly a lot more efficient.</p>",
          "rawMarkdown": "I've actually thought about building a 4x RTX 8000 server but I'm waiting to see retail availability and pricing of the A100 PCIe component. \n\nAbout not being able to fit the Toxic Comment models in memory, I was referring more to pretraining MLM models. They have a much bigger memory footprint because of their classifier head (shape: seq len, vocab size). I'm going to try to use the electra approach in the future to pretrain - it's supposedly a lot more efficient.",
          "votes": 1
        },
        {
          "id": 904123,
          "postDate": "2020-06-27T11:33:34.127Z",
          "content": "<p>I tried ELECTRA pretrained model in Tweet Sentiment, it wasn't anywhere near RoBERTa.  I wonder if it is because of the training data they used in pre training or if ELECTRA method isn't that good.</p>",
          "rawMarkdown": "I tried ELECTRA pretrained model in Tweet Sentiment, it wasn't anywhere near RoBERTa.  I wonder if it is because of the training data they used in pre training or if ELECTRA method isn't that good."
        }
      ]
    },
    {
      "id": 895426,
      "postDate": "2020-06-21T10:38:01.673Z",
      "content": "<p>jigsaw1: we can use CPU to win.\njigsaw2: we can use free kaggle GPU(no time limit).\njigsaw3: we only have 30 hours per week.</p>",
      "rawMarkdown": "jigsaw1: we can use CPU to win.\njigsaw2: we can use free kaggle GPU(no time limit).\njigsaw3: we only have 30 hours per week.",
      "votes": 8,
      "replies": [
        {
          "id": 895876,
          "postDate": "2020-06-21T17:03:32.303Z",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> , Good points.</p>",
          "rawMarkdown": "@mcggood , Good points."
        }
      ]
    },
    {
      "id": 895406,
      "postDate": "2020-06-21T10:18:10.823Z",
      "content": "<p>The best models in this competition can only be fitted with TPUs or really good local hardware. TPUs are really shaky, and not many have good HW at home. Also Tweet deadline extension.</p>\n\n<p>I think those are the reasons. </p>\n\n<p>On the other hand, there is no hidden private data and subbing public kernels is straight-forward. That usually increases participation rates.</p>\n\n<p>But I agree, participant numbers seem to be going down lately.</p>",
      "rawMarkdown": "The best models in this competition can only be fitted with TPUs or really good local hardware. TPUs are really shaky, and not many have good HW at home. Also Tweet deadline extension.\n\nI think those are the reasons. \n\nOn the other hand, there is no hidden private data and subbing public kernels is straight-forward. That usually increases participation rates.\n\nBut I agree, participant numbers seem to be going down lately.",
      "votes": 5,
      "replies": [
        {
          "id": 895896,
          "postDate": "2020-06-21T17:11:15.107Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>, I agree hence my motivation to make this post. I did not want to mention the TPU/GPU restriction as I have argued against it so much when it was 1st implemented. I am glad you and a couple of others mentioned it.</p>",
          "rawMarkdown": "@philippsinger, I agree hence my motivation to make this post. I did not want to mention the TPU/GPU restriction as I have argued against it so much when it was 1st implemented. I am glad you and a couple of others mentioned it."
        }
      ]
    },
    {
      "id": 895396,
      "postDate": "2020-06-21T10:06:26.140Z",
      "content": "<p>Even <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge\">deepfake-detection-challenge</a> only had 2k teams, with 1 million USD in prizes! \nApart from M5, no other comp at the moment has much more than 1k teams, perhaps Kaggle is losing a bit of its popularity...</p>",
      "rawMarkdown": "Even [deepfake-detection-challenge](https://www.kaggle.com/c/deepfake-detection-challenge) only had 2k teams, with 1 million USD in prizes! \nApart from M5, no other comp at the moment has much more than 1k teams, perhaps Kaggle is losing a bit of its popularity...",
      "votes": 5,
      "replies": [
        {
          "id": 895898,
          "postDate": "2020-06-21T17:12:27.587Z",
          "content": "<p><a href=\"/hmendonca\">@hmendonca</a>, you may be right Kaggle is losing its touch.</p>",
          "rawMarkdown": "@hmendonca, you may be right Kaggle is losing its touch."
        },
        {
          "id": 895995,
          "postDate": "2020-06-21T18:33:26.747Z",
          "content": "<p>It was a 500gb download to even get started on that competition. I think that was a stumbling block for many</p>",
          "rawMarkdown": "It was a 500gb download to even get started on that competition. I think that was a stumbling block for many",
          "votes": 2
        },
        {
          "id": 896136,
          "postDate": "2020-06-21T22:36:33.613Z",
          "content": "<p><a href=\"/ryches\">@ryches</a>, that was a no go for me as well just for that reason.</p>",
          "rawMarkdown": "@ryches, that was a no go for me as well just for that reason."
        }
      ]
    },
    {
      "id": 897540,
      "postDate": "2020-06-23T00:23:44.250Z",
      "content": "<p>They should start using constrained Environments IMO now! And Kaggle is losing it's charm as well;</p>",
      "rawMarkdown": "They should start using constrained Environments IMO now! And Kaggle is losing it's charm as well;",
      "votes": 1,
      "replies": [
        {
          "id": 898997,
          "postDate": "2020-06-23T21:59:37.387Z",
          "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a>, yes Kaggle started losing its charm as soon as they restricted the GPU use to 30hrs/week.</p>",
          "rawMarkdown": "@adityaecdrid, yes Kaggle started losing its charm as soon as they restricted the GPU use to 30hrs/week."
        }
      ]
    },
    {
      "id": 895669,
      "postDate": "2020-06-21T14:27:50.257Z",
      "content": "<p>A lot of factors are contributing to low team numbers: Covid causes lack of motivation, loss of hardware for some because they couldn't go to work, loss of job completely, etc. Not only that but recent competitions have also either have no prizes, unique and difficult problems, or requires a lot of hardware to compete in.</p>\n\n<p>For this comp specifically: its not a really unique or interesting problem its the third one, tweet comp going at the same time which is much more interesting and it ended first, and hardware needs.</p>",
      "rawMarkdown": "A lot of factors are contributing to low team numbers: Covid causes lack of motivation, loss of hardware for some because they couldn't go to work, loss of job completely, etc. Not only that but recent competitions have also either have no prizes, unique and difficult problems, or requires a lot of hardware to compete in.\n\nFor this comp specifically: its not a really unique or interesting problem its the third one, tweet comp going at the same time which is much more interesting and it ended first, and hardware needs.",
      "votes": 1,
      "replies": [
        {
          "id": 895889,
          "postDate": "2020-06-21T17:07:34.020Z",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>, I agree Tweet is much more interesting hence I joined that first and was ready to pass on this one. I only joined Jigsaw when I realised I will have some free time after Tweet a few days before the entry dateline.</p>",
          "rawMarkdown": "@greatgamedota, I agree Tweet is much more interesting hence I joined that first and was ready to pass on this one. I only joined Jigsaw when I realised I will have some free time after Tweet a few days before the entry dateline."
        }
      ]
    },
    {
      "id": 895040,
      "postDate": "2020-06-21T04:10:12.733Z",
      "content": "<p>Average number of participants goes down each year. The average for this year is around 1500 participants per competition. In the past it was almost twice more. </p>",
      "rawMarkdown": "Average number of participants goes down each year. The average for this year is around 1500 participants per competition. In the past it was almost twice more. ",
      "votes": 1,
      "replies": [
        {
          "id": 895054,
          "postDate": "2020-06-21T04:26:00.623Z",
          "content": "<p><a href=\"/muhakabartay\">@muhakabartay</a> interesting, I did not pay much attention to the number of participants recently but I expected it to be higher than last year because of the covid19 lockdown.</p>",
          "rawMarkdown": "@muhakabartay interesting, I did not pay much attention to the number of participants recently but I expected it to be higher than last year because of the covid19 lockdown."
        }
      ]
    },
    {
      "id": 894947,
      "postDate": "2020-06-20T23:55:58.330Z",
      "content": "<p>I suppose it is because there was also the \"sentiment extraction\"  almost at the same time</p>",
      "rawMarkdown": "I suppose it is because there was also the \"sentiment extraction\"  almost at the same time",
      "votes": 2,
      "replies": [
        {
          "id": 894952,
          "postDate": "2020-06-21T00:18:53.227Z",
          "content": "<p><a href=\"/ludovick\">@ludovick</a>, that is definately a factor as I came from that competition and almost didn't enter this one because of Tweet. However, I was suprised when I checked last year's compared to their 1st competition.</p>",
          "rawMarkdown": "@ludovick, that is definately a factor as I came from that competition and almost didn't enter this one because of Tweet. However, I was suprised when I checked last year's compared to their 1st competition."
        }
      ]
    },
    {
      "id": 902906,
      "postDate": "2020-06-26T12:45:29.917Z",
      "content": "<p>Thanks all who upvoted, commented here and making it a lively discussion. I have updated my main post with a summary of the most voted points relating to the topic at hand.</p>",
      "rawMarkdown": "Thanks all who upvoted, commented here and making it a lively discussion. I have updated my main post with a summary of the most voted points relating to the topic at hand."
    },
    {
      "id": 900440,
      "postDate": "2020-06-24T20:00:40.110Z",
      "content": "<p>There are lots of other competitions running concurrently.  Was it the case 2 years ago?</p>",
      "rawMarkdown": "There are lots of other competitions running concurrently.  Was it the case 2 years ago?",
      "replies": [
        {
          "id": 902879,
          "postDate": "2020-06-26T12:21:07.273Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, I just checked and from what I can see 2yrs ago Toxic was the only competition that ended in March. That year's Santa which I also took part in ended mid January and IEEE finished on February 8th. So perhaps it is the combination of other competition running at the same time which in the case of Tweet also an NLP one, coupled with the limit on GPU usage to 30hrs per week.</p>",
          "rawMarkdown": "@cpmpml, I just checked and from what I can see 2yrs ago Toxic was the only competition that ended in March. That year's Santa which I also took part in ended mid January and IEEE finished on February 8th. So perhaps it is the combination of other competition running at the same time which in the case of Tweet also an NLP one, coupled with the limit on GPU usage to 30hrs per week.",
          "votes": 1
        }
      ]
    },
    {
      "id": 896147,
      "postDate": "2020-06-21T23:03:27.177Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 895882,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-06-21T17:05:17.233000",
      "content": "<p>It's getting way less accessible. Needing to use TPU's to even have the ability to train these massive models is a large barrier. It was easy when it was TF-IDF and LSTM. I feel like the compute requirement, complexity and difficulty seems to keep going up. It is not very easy for a newcomer to jump in from zero on most of these competitions. </p>",
      "votes": 8,
      "replies": [
        {
          "id": 895905,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-06-21T17:19:24.327000",
          "content": "<p>My first competition was Quora, and to me it was nearly the perfect setup, except of the two stage re-runs which they would not need to do now anylonger having synchronous kernels.</p>\n\n<p>It had limited GPU time for everyone, no external data except pre-provided vectors, people developing elegant simple models, and in the end nice solutions.</p>\n\n<p>Nowadays it is getting more complex to run the models, external data matters a lot often, and everything is getting fuzzy. I personally, really like constrained environments, even though that many people complain about kernel competitions.</p>",
          "votes": 15,
          "replies": []
        },
        {
          "id": 895926,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T17:34:45.777000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>, Quora was my 2nd competition and I was new to NLP as well found that Quora contest  as a fartile ground for learning and experimentation.</p>\n\n<p>UPDATE:</p>\n\n<p>I did not realize you meant the 2nd Quora competition, my post above is refering to the <a href=\"https://www.kaggle.com/c/quora-question-pairs\">Quora Question Pairs</a> competition 😄 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 895931,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-06-21T17:37:34.857000",
          "content": "<p>I agree that having a constrained environment is nice, but sometimes the creativity that removing those restrictions provides is nice too. In that Quora competition if we were allowed to use whatever we wanted then BERT would have dominated there which makes you wonder if that was really the best solution for the host. It is a bit of an arms race unless someone puts a cap on our solutions. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 895936,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-06-21T17:41:27.667000",
          "content": "<p>I really like constrained environments too (This may seem biased given my two gold medals come from Kernel competitions ). Even though Mercari was a bit too constrained IMHO.</p>\n\n<p>But this competition helped me to get familiar with TPUs, and I have to  say I really envoy now training quickly on large text or Images dataset with TPUs. I will be using 100% TPU on SIIM ISIC competition</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 895950,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-06-21T17:50:47.113000",
          "content": "<p><a href=\"/ryches\">@ryches</a> The question is what is more valuable? Finetuning Bert or finding simpler models that are on-par with Bert? I think both have its value.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 895994,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-06-21T18:32:44.600000",
          "content": "<p>It depends on the the value of a more accurate model. I would argue that scraping the last % out of this toxic comment stuff probably isn't really that valuable and a more constrained solution is probably what they want. That is common of a lot of kaggle problems though.</p>\n\n<p>Definitely value to exploring both constrained and unconstrained approaches</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 896025,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T18:59:08.193000",
          "content": "<p>The way I see it is that developing a simple high performing model in a constrained environment challenges me to dig deeper. This is why I prefer that than fine tuning the new kid on the block like BERT.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 896060,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-06-21T19:41:12.863000",
          "content": "<p>However, BERT is easier to parallelize than LSTM...And with the same amount of data , you can  make it run faster than LSTM (by choosing for instance which hidden layers to freeze/fine-tune) while still outperforming LSTM. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 896062,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-06-21T19:43:54.477000",
          "content": "<p>My issue is that noone ever tries simpler architectures nowadays because you have a bias towards starting with SOTA models. Constrained Kaggle competitions might force you to try other stuff outside of current research, and maybe it would lead to a higher chance of finding new architectures or solutions to the problem at hand.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 896077,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T20:02:32.707000",
          "content": "<p><a href=\"/serigne\">@serigne</a>, agree and I am not against BERT by the way. Various fine tuned and nicely designed RoBERTa models got my team out of trouble from #241 to 41st in the Tweet competition.</p>\n\n<p><a href=\"/philippsinger\">@philippsinger</a> my sentiments exactly. </p>\n\n<p>I just realised the content of my post must be really really boring going by the number of upvotes😃 Having said that, I am so happy it pompted all these opinions I am reading. Thanks guys and girrls.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 896205,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-22T00:53:22.890000",
          "content": "<p>The problem with artificial constraints is you often end up emphasizing engineering (e.g., working around CPU/memory bottlenecks or runtime limits). That's a useful skillset but not more important than being able to tweak SOTA models. </p>\n\n<p>With academia moving to using TPU pods, model sizes are going to increase. There's also research showing that given a fixed compute budget, it's better to train large models for a short period and compressing them, rather than training small/moderately-sized models for longer. Not a bad idea for a Kaggler to acclimatize themselves to handling large models. </p>\n\n<p>RL and generative problems (e.g., Halite, ARC) are what you're looking for if you want something requiring more creativity. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 896210,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-22T01:05:52.297000",
          "content": "<p><a href=\"/leecming\">@leecming</a> you are quite right about ARC, I was too busy at work to be able to participate at the time but believe that is a glimse into the future of ML. Halite I thought look very interesting from the Kaggle email invitation, I will be taking a look at that after this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 896658,
          "author_name": "godelscat",
          "author_url": "",
          "post_date": "2020-06-22T10:54:59.927000",
          "content": "<p>Yeah, I totally agree. I'm a newbie to nlp, and this is my first nlp competition. In fact, I just started to learn nlp months ago by cs224n. I spent a lot of time trying to run these super large models on TPUs with pytorch at the beginning. This is painful, but learnt a lot.  Thanks for those excellent public kernels.\nBWT, may I ask what the general parameters you try first while dealing with these large models? Is it like \"lr: 5e-5, 3e-5, 2e-5; epochs: 2, 3, 4\" as suggested by the BERT paper? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 900208,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-24T17:24:21.393000",
          "content": "<p><a href=\"/godelscat\">@godelscat</a>, you can try going through some of the starter BERT model kernels. The steps in the following articles is a good one to follow.</p>\n\n<ol>\n<li><a href=\"https://towardsdatascience.com/fine-tuning-bert-for-text-classification-with-farm-2880665065e2\">https://towardsdatascience.com/fine-tuning-bert-for-text-classification-with-farm-2880665065e2</a></li>\n<li><a href=\"https://mccormickml.com/2019/07/22/BERT-fine-tuning/\">https://mccormickml.com/2019/07/22/BERT-fine-tuning/</a></li>\n<li><a href=\"https://medium.com/curation-corporation/fine-tuning-bart-for-abstractive-text-summarisation-with-fastai2-d7a2ad676a13\">https://medium.com/curation-corporation/fine-tuning-bart-for-abstractive-text-summarisation-with-fastai2-d7a2ad676a13</a></li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 900450,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-06-24T20:02:59.217000",
          "content": "<blockquote>\n  <p>My issue is that noone ever tries simpler architectures nowadays</p>\n</blockquote>\n\n<p>I used linear regression in Covid forecasting, and this was competitive with best models once we factor out difference in training data time span.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 900601,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-24T23:32:30.943000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, yep that is the thing. In fact I trained FM-FTRL, NBSVM, Single-Wordbatch-LGBM, and PooledGRU-Word2Vec-using-Gensim models in this competition that helped my ensemble score. These models scored low in the range of 0.8750 to 0.9000 on the public LB. I did not know what some of them scored until after the competition in late submission because I was submitting the ensemble based on CV.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 901462,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-06-25T13:53:23.990000",
          "content": "<blockquote>\n  <p>With academia moving to using TPU pods, , model sizes are going to increase.</p>\n</blockquote>\n\n<p>Is TPU the reason really?</p>\n\n<p>AFAIK, the largest model ever was trained on GPU: <a href=\"https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/\">https://www.microsoft.com/en-us/research/blog/turing-nlg-a-17-billion-parameter-language-model-by-microsoft/</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 902075,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-25T23:03:17.413000",
          "content": "<p>GPU too - Nvidia recently announced their A100 GPU with 40GB VRAM. There's an old saying for hard drives about how applications increase in size to fit growing hard drive capacities. Same with VRAM. </p>\n\n<p>I have two RTX Titans at home with a combined 48GB VRAM. Even with gradient accumulation and mixed precision training, I'm having trouble squeezing SOTA NLP models for pretraining tasks e.g. MLM.</p>\n\n<p>I've half thought about buying a DGX for home use. Can you get me an employee discount :P</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 902760,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-06-26T10:45:27.900000",
          "content": "<p>There are no employee discounts for DGX machines.  I guess only few employees could afford these.  And except for DGX Stations, DGX machines are mean to be used in racks in datacenters.  </p>\n\n<p>We can get PCIe GPU cards eg 2080 Ti,  with a small discount, but that's not what you are discussing, right?  And of course we can't resell.</p>\n\n<p>An alternative to V100 and DGX Station that may be worth considering is RTX 8000 cards.  They are PCIe cards, with 48 GB memory. Jason Antic, the deoldify guy just moved to a 4x RTX 8000 machine for instance: <a href=\"https://twitter.com/citnaj/status/1271892395798392832\">https://twitter.com/citnaj/status/1271892395798392832</a></p>\n\n<p>I am puzzled by your comment about SOTA models.  I trained my models on one V100 at a time.  It has 32GB memory, which is less than what you have.  I'm forced to us small batch size, and as you said, gradient accumulation and mixed precision.  But I can train any of the huggingfaces model.</p>\n\n<p>What may be worrisome is if people take advantage of A100 expanded capacity to move SOTA models beyond what V100 can handle.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 902824,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2020-06-26T11:39:48.637000",
          "content": "<p>@CPMP but the price of 4 RTX Quadro 8000 begin to be a bit insane right ? 😄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 903626,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2020-06-27T02:10:45.883000",
          "content": "<p>I've actually thought about building a 4x RTX 8000 server but I'm waiting to see retail availability and pricing of the A100 PCIe component. </p>\n\n<p>About not being able to fit the Toxic Comment models in memory, I was referring more to pretraining MLM models. They have a much bigger memory footprint because of their classifier head (shape: seq len, vocab size). I'm going to try to use the electra approach in the future to pretrain - it's supposedly a lot more efficient.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 904123,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-06-27T11:33:34.127000",
          "content": "<p>I tried ELECTRA pretrained model in Tweet Sentiment, it wasn't anywhere near RoBERTa.  I wonder if it is because of the training data they used in pre training or if ELECTRA method isn't that good.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 895426,
      "author_name": "MaChaogong",
      "author_url": "",
      "post_date": "2020-06-21T10:38:01.673000",
      "content": "<p>jigsaw1: we can use CPU to win.\njigsaw2: we can use free kaggle GPU(no time limit).\njigsaw3: we only have 30 hours per week.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 895876,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T17:03:32.303000",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> , Good points.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 895406,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-06-21T10:18:10.823000",
      "content": "<p>The best models in this competition can only be fitted with TPUs or really good local hardware. TPUs are really shaky, and not many have good HW at home. Also Tweet deadline extension.</p>\n\n<p>I think those are the reasons. </p>\n\n<p>On the other hand, there is no hidden private data and subbing public kernels is straight-forward. That usually increases participation rates.</p>\n\n<p>But I agree, participant numbers seem to be going down lately.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 895896,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T17:11:15.107000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>, I agree hence my motivation to make this post. I did not want to mention the TPU/GPU restriction as I have argued against it so much when it was 1st implemented. I am glad you and a couple of others mentioned it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 895396,
      "author_name": "Henrique Mendonça",
      "author_url": "",
      "post_date": "2020-06-21T10:06:26.140000",
      "content": "<p>Even <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge\">deepfake-detection-challenge</a> only had 2k teams, with 1 million USD in prizes! \nApart from M5, no other comp at the moment has much more than 1k teams, perhaps Kaggle is losing a bit of its popularity...</p>",
      "votes": 5,
      "replies": [
        {
          "id": 895898,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T17:12:27.587000",
          "content": "<p><a href=\"/hmendonca\">@hmendonca</a>, you may be right Kaggle is losing its touch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 895995,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-06-21T18:33:26.747000",
          "content": "<p>It was a 500gb download to even get started on that competition. I think that was a stumbling block for many</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 896136,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T22:36:33.613000",
          "content": "<p><a href=\"/ryches\">@ryches</a>, that was a no go for me as well just for that reason.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 897540,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-06-23T00:23:44.250000",
      "content": "<p>They should start using constrained Environments IMO now! And Kaggle is losing it's charm as well;</p>",
      "votes": 1,
      "replies": [
        {
          "id": 898997,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-23T21:59:37.387000",
          "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a>, yes Kaggle started losing its charm as soon as they restricted the GPU use to 30hrs/week.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 895669,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-06-21T14:27:50.257000",
      "content": "<p>A lot of factors are contributing to low team numbers: Covid causes lack of motivation, loss of hardware for some because they couldn't go to work, loss of job completely, etc. Not only that but recent competitions have also either have no prizes, unique and difficult problems, or requires a lot of hardware to compete in.</p>\n\n<p>For this comp specifically: its not a really unique or interesting problem its the third one, tweet comp going at the same time which is much more interesting and it ended first, and hardware needs.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 895889,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T17:07:34.020000",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>, I agree Tweet is much more interesting hence I joined that first and was ready to pass on this one. I only joined Jigsaw when I realised I will have some free time after Tweet a few days before the entry dateline.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 895040,
      "author_name": "Mukharbek Organokov",
      "author_url": "",
      "post_date": "2020-06-21T04:10:12.733000",
      "content": "<p>Average number of participants goes down each year. The average for this year is around 1500 participants per competition. In the past it was almost twice more. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 895054,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T04:26:00.623000",
          "content": "<p><a href=\"/muhakabartay\">@muhakabartay</a> interesting, I did not pay much attention to the number of participants recently but I expected it to be higher than last year because of the covid19 lockdown.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 894947,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2020-06-20T23:55:58.330000",
      "content": "<p>I suppose it is because there was also the \"sentiment extraction\"  almost at the same time</p>",
      "votes": 2,
      "replies": [
        {
          "id": 894952,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-21T00:18:53.227000",
          "content": "<p><a href=\"/ludovick\">@ludovick</a>, that is definately a factor as I came from that competition and almost didn't enter this one because of Tweet. However, I was suprised when I checked last year's compared to their 1st competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 902906,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2020-06-26T12:45:29.917000",
      "content": "<p>Thanks all who upvoted, commented here and making it a lively discussion. I have updated my main post with a summary of the most voted points relating to the topic at hand.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 900440,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-06-24T20:00:40.110000",
      "content": "<p>There are lots of other competitions running concurrently.  Was it the case 2 years ago?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 902879,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-06-26T12:21:07.273000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a>, I just checked and from what I can see 2yrs ago Toxic was the only competition that ended in March. That year's Santa which I also took part in ended mid January and IEEE finished on February 8th. So perhaps it is the combination of other competition running at the same time which in the case of Tweet also an NLP one, coupled with the limit on GPU usage to 30hrs per week.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 896147,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-21T23:03:27.177000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "894877": "I noticed that the number of teams in this comeptition is a lot lower than last year's which prompted me to check all the numbers. I thought it is interesting that the number of team decrease is roughly the same yearly plus some decay factor😃 . See below:-\n\n1. [Toxic Comment Classification Challenge](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge) - total participation is 4,550 teams.\n\n2. [Jigsaw Unintended Bias in Toxicity Classification](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification) - total participation is 3,165 teams (down by 1385)\n\n3. [Jigsaw Multilingual Toxic Comment Classification](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/leaderboard) - total participation is 1,480 teams (down by 1685). This one may go down on closing after cheaters removal.\n\n### UPDATE:-\n\nI am updating my main post with a summary of the most voted comments for anyone who did not want to go through them.\n\n1. Another very engaging and arguably even more interesting NLP competition going on at the same time i.e. [Tweet Sentiment Extraction](https://www.kaggle.com/c/tweet-sentiment-extraction).\n2. In the past we had unlimited GPU time for everyone, no external data except pre-provided vectors, people developing elegant simple models, and in the end nice solutions. @philippsinger \n3. - jigsaw1: we can use CPU to win. - jigsaw2: we can use free kaggle GPU(no time limit). - jigsaw3: we only have 30 hours per week @mcggood \n4. Needing to use TPU's to even have the ability to train these massive models is a large barrier. It was easy when it was TF-IDF and LSTM. Complexity and difficulty seems to keep going up. It is not very easy for a newcomer to jump in from zero on most of these competitions. @ryches \n",
    "895882": "It's getting way less accessible. Needing to use TPU's to even have the ability to train these massive models is a large barrier. It was easy when it was TF-IDF and LSTM. I feel like the compute requirement, complexity and difficulty seems to keep going up. It is not very easy for a newcomer to jump in from zero on most of these competitions. ",
    "895426": "jigsaw1: we can use CPU to win.\njigsaw2: we can use free kaggle GPU(no time limit).\njigsaw3: we only have 30 hours per week.",
    "895406": "The best models in this competition can only be fitted with TPUs or really good local hardware. TPUs are really shaky, and not many have good HW at home. Also Tweet deadline extension.\n\nI think those are the reasons. \n\nOn the other hand, there is no hidden private data and subbing public kernels is straight-forward. That usually increases participation rates.\n\nBut I agree, participant numbers seem to be going down lately.",
    "895396": "Even [deepfake-detection-challenge](https://www.kaggle.com/c/deepfake-detection-challenge) only had 2k teams, with 1 million USD in prizes! \nApart from M5, no other comp at the moment has much more than 1k teams, perhaps Kaggle is losing a bit of its popularity...",
    "897540": "They should start using constrained Environments IMO now! And Kaggle is losing it's charm as well;",
    "895669": "A lot of factors are contributing to low team numbers: Covid causes lack of motivation, loss of hardware for some because they couldn't go to work, loss of job completely, etc. Not only that but recent competitions have also either have no prizes, unique and difficult problems, or requires a lot of hardware to compete in.\n\nFor this comp specifically: its not a really unique or interesting problem its the third one, tweet comp going at the same time which is much more interesting and it ended first, and hardware needs.",
    "895040": "Average number of participants goes down each year. The average for this year is around 1500 participants per competition. In the past it was almost twice more. ",
    "894947": "I suppose it is because there was also the \"sentiment extraction\"  almost at the same time",
    "902906": "Thanks all who upvoted, commented here and making it a lively discussion. I have updated my main post with a summary of the most voted points relating to the topic at hand.",
    "900440": "There are lots of other competitions running concurrently.  Was it the case 2 years ago?",
    "896147": ""
  }
}