{
  "id": 154908,
  "title": "OpenAI's new GPT-3 Model",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/154908",
  "author_name": "",
  "post_date": "2020-05-30T11:59:00.672489200Z",
  "votes": 16,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Just <a href=\"https://arxiv.org/abs/2005.14165\">released</a>  .. the task agnostic, zero / one shot sounds great, the only problem is 175B parameters (~ 350 G GPU memory with 16 bit?) 😄 😄 </p>",
  "messages": [
    {
      "id": "867543",
      "postDate": "05/30/2020 11:59:00",
      "content": "<p>Just <a href=\"https://arxiv.org/abs/2005.14165\">released</a>  .. the task agnostic, zero / one shot sounds great, the only problem is 175B parameters (~ 350 G GPU memory with 16 bit?) 😄 😄 </p>",
      "rawMarkdown": "Just [released](https://arxiv.org/abs/2005.14165)  .. the task agnostic, zero / one shot sounds great, the only problem is 175B parameters (~ 350 G GPU memory with 16 bit?) 😄 😄",
      "votes": null
    },
    {
      "id": "867571",
      "postDate": "05/30/2020 12:18:30",
      "content": "<p>Wow! so XLM-R is just too small! \nIt looks like, that the secret of humans learning from just a few samples is just because of our large brains? ;)\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F6bfb4878af94fc632eb0a98a19e3e9fe%2Fgpt3.png?generation=1590841027905156&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Wow! so XLM-R is just too small! \nIt looks like, that the secret of humans learning from just a few samples is just because of our large brains? ;)\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F6bfb4878af94fc632eb0a98a19e3e9fe%2Fgpt3.png?generation=1590841027905156&amp;alt=media)",
      "votes": null
    },
    {
      "id": "867989",
      "postDate": "05/30/2020 19:56:05",
      "content": "<p>GPT3 will be like :D </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F0e61c7c70114249f40754b09629b3d24%2Fimages%20-%202020-05-31T012413.281.jpeg?generation=1590868527705882&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "GPT3 will be like :D \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F0e61c7c70114249f40754b09629b3d24%2Fimages%20-%202020-05-31T012413.281.jpeg?generation=1590868527705882&amp;alt=media)",
      "votes": null
    },
    {
      "id": "868217",
      "postDate": "05/31/2020 04:11:25",
      "content": "<p>Any tutorial?</p>",
      "rawMarkdown": "Any tutorial?",
      "votes": null
    },
    {
      "id": "868298",
      "postDate": "05/31/2020 06:00:32",
      "content": "<p>Gigantic models getting much bigger. We need to think of novel approach to finetune these models on less resources.</p>",
      "rawMarkdown": "Gigantic models getting much bigger. We need to think of novel approach to finetune these models on less resources.",
      "votes": null
    },
    {
      "id": "868370",
      "postDate": "05/31/2020 06:59:00",
      "content": "<p>They haven't released the source code yet, just the pre print paper a few days back that I shared above. </p>",
      "rawMarkdown": "They haven't released the source code yet, just the pre print paper a few days back that I shared above.",
      "votes": null
    },
    {
      "id": "868379",
      "postDate": "05/31/2020 07:05:29",
      "content": "<p>I said this before and I will 100 percent say it now.</p>\n\n<blockquote>\n  <p>Before it was about an elegant approach and well-written code. Now it is basically \"MY MODEL TRAINED FOR 1 WEEK ON 16 GPUS WITH 500B PARAMETERS. BEAT THAT GOOGLE!\"</p>\n</blockquote>",
      "rawMarkdown": "I said this before and I will 100 percent say it now.\n\n&gt; Before it was about an elegant approach and well-written code. Now it is basically \"MY MODEL TRAINED FOR 1 WEEK ON 16 GPUS WITH 500B PARAMETERS. BEAT THAT GOOGLE!\"",
      "votes": null
    },
    {
      "id": "868845",
      "postDate": "05/31/2020 14:24:51",
      "content": "<p>OpenAI = ClosedAI</p>\n\n<p>They will likely release nothing for a while because they think \"it may be too dangerous\" lol. </p>\n\n<p>And Given the size of the model , forget about using it here ( Unless there are smaller versions)</p>",
      "rawMarkdown": "OpenAI = ClosedAI\n\nThey will likely release nothing for a while because they think \"it may be too dangerous\" lol. \n\nAnd Given the size of the model , forget about using it here ( Unless there are smaller versions)",
      "votes": null
    },
    {
      "id": "868872",
      "postDate": "05/31/2020 14:45:42",
      "content": "<p>Language modelling is difficult, particularly predicting new task. </p>\n\n<p>So , they think the only way to overcome this, is to feed the model with as much data as possible  (That's not what I'd call \"Intelligence\" :p) </p>\n\n<p>Anyway Google has started a smart move with the Electra Model. </p>\n\n<p>But as you say, there should be room for elegant NLP papers instead of just focusing on these SOTA performance. </p>\n\n<p>Attention is a cutting Edge feature compared to the really slow and unparallelizable LSTM and both the Transformer and Bert papers were big game changers. But may be it's time to move into \"smarter\" directions than these \" Bigger and Bigger\" models that no one or very few will use in real life Applications. </p>",
      "rawMarkdown": "Language modelling is difficult, particularly predicting new task. \n\nSo , they think the only way to overcome this, is to feed the model with as much data as possible  (That's not what I'd call \"Intelligence\" :p) \n\nAnyway Google has started a smart move with the Electra Model. \n\nBut as you say, there should be room for elegant NLP papers instead of just focusing on these SOTA performance. \n\nAttention is a cutting Edge feature compared to the really slow and unparallelizable LSTM and both the Transformer and Bert papers were big game changers. But may be it's time to move into \"smarter\" directions than these \" Bigger and Bigger\" models that no one or very few will use in real life Applications.",
      "votes": null
    },
    {
      "id": "868930",
      "postDate": "05/31/2020 15:16:42",
      "content": "<blockquote>\n  <p>Attention is a cutting Edge feature compared to the really slow and unparallelizable LSTM and both the Transformers and Bert were big game changers. But may be it's time to move into \"smarter\" directions than these \" Bigger and Bigger\" models that no one or very few will use in real life Applications.</p>\n</blockquote>\n\n<p>I feel that we should try to look at the next big breakthrough in NLP/CV in general rather than wasting our time on these titanic models. </p>\n\n<p>I would like to see something novel like Attention crop up, or maybe the Transformer.... anything that is \"novel\". Not \"bigger\".</p>",
      "rawMarkdown": "&gt; Attention is a cutting Edge feature compared to the really slow and unparallelizable LSTM and both the Transformers and Bert were big game changers. But may be it's time to move into \"smarter\" directions than these \" Bigger and Bigger\" models that no one or very few will use in real life Applications.\n\nI feel that we should try to look at the next big breakthrough in NLP/CV in general rather than wasting our time on these titanic models. \n\nI would like to see something novel like Attention crop up, or maybe the Transformer.... anything that is \"novel\". Not \"bigger\".",
      "votes": null
    },
    {
      "id": "868949",
      "postDate": "05/31/2020 15:38:34",
      "content": "<p>So true hahaha.</p>",
      "rawMarkdown": "So true hahaha.",
      "votes": null
    },
    {
      "id": "869065",
      "postDate": "05/31/2020 17:53:58",
      "content": "<p>175B params - thats crazy. i will leave it here just in case someone wants to follow that <a href=\"https://github.com/openai/gpt-3\">https://github.com/openai/gpt-3</a></p>",
      "rawMarkdown": "175B params - thats crazy. i will leave it here just in case someone wants to follow that [https://github.com/openai/gpt-3](https://github.com/openai/gpt-3)",
      "votes": null
    },
    {
      "id": "872992",
      "postDate": "06/03/2020 17:12:50",
      "content": "<p>I think that's really true. I mean even if we assume a single weight to be a single neuron we're still not nearly as big as the brain. Maybe the answer is infact bigger networks.</p>",
      "rawMarkdown": "I think that's really true. I mean even if we assume a single weight to be a single neuron we're still not nearly as big as the brain. Maybe the answer is infact bigger networks.",
      "votes": null
    },
    {
      "id": "877743",
      "postDate": "06/07/2020 23:08:21",
      "content": "<p><a href=\"https://gateway.on24.com/wcc/eh/2336732/lp/2357495/burning-the-world-with-machine-learning%3A-understanding-the-carbon-footprint-of-deep-learning/\">https://gateway.on24.com/wcc/eh/2336732/lp/2357495/burning-the-world-with-machine-learning%3A-understanding-the-carbon-footprint-of-deep-learning/</a></p>",
      "rawMarkdown": "https://gateway.on24.com/wcc/eh/2336732/lp/2357495/burning-the-world-with-machine-learning%3A-understanding-the-carbon-footprint-of-deep-learning/",
      "votes": null
    },
    {
      "id": "883670",
      "postDate": "06/12/2020 20:18:07",
      "content": "<p>I think people who're designing such a gigantic network should be constrained Kaggle like computing resources, like 30 hours GPU and TPU quote for their research. So, they have to think to design reasonalbe size awesome network 😅 </p>",
      "rawMarkdown": "I think people who're designing such a gigantic network should be constrained Kaggle like computing resources, like 30 hours GPU and TPU quote for their research. So, they have to think to design reasonalbe size awesome network 😅",
      "votes": null
    },
    {
      "id": "992728",
      "postDate": "08/31/2020 11:51:57",
      "content": "<p>With proper API keys, could one access APIs straight from the Kaggle notebooks?</p>",
      "rawMarkdown": "With proper API keys, could one access APIs straight from the Kaggle notebooks?",
      "votes": null
    },
    {
      "id": "2204501",
      "postDate": "03/31/2023 16:09:13",
      "content": "<p>So i made an api wrapper for the Da Vinci model provided by open-ai you can easily run it and use it in web apps and android apps. Please upvote if you like it :D.</p>",
      "rawMarkdown": "So i made an api wrapper for the Da Vinci model provided by open-ai you can easily run it and use it in web apps and android apps. Please upvote if you like it :D.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 992728,
      "author_name": "gabrielberzescu",
      "author_url": "",
      "post_date": "08/31/2020 11:51:57",
      "content": "<p>With proper API keys, could one access APIs straight from the Kaggle notebooks?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2204501,
      "author_name": "amanalisiddiqui",
      "author_url": "",
      "post_date": "03/31/2023 16:09:13",
      "content": "<p>So i made an api wrapper for the Da Vinci model provided by open-ai you can easily run it and use it in web apps and android apps. Please upvote if you like it :D.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 867571,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "05/30/2020 12:18:30",
      "content": "<p>Wow! so XLM-R is just too small! \nIt looks like, that the secret of humans learning from just a few samples is just because of our large brains? ;)\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F6bfb4878af94fc632eb0a98a19e3e9fe%2Fgpt3.png?generation=1590841027905156&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 867989,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "05/30/2020 19:56:05",
          "content": "<p>GPT3 will be like :D </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F0e61c7c70114249f40754b09629b3d24%2Fimages%20-%202020-05-31T012413.281.jpeg?generation=1590868527705882&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 872992,
          "author_name": "jonykarki",
          "author_url": "",
          "post_date": "06/03/2020 17:12:50",
          "content": "<p>I think that's really true. I mean even if we assume a single weight to be a single neuron we're still not nearly as big as the brain. Maybe the answer is infact bigger networks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 868217,
      "author_name": "anuraglahon",
      "author_url": "",
      "post_date": "05/31/2020 04:11:25",
      "content": "<p>Any tutorial?</p>",
      "votes": null,
      "replies": [
        {
          "id": 868370,
          "author_name": "moizsaifee",
          "author_url": "",
          "post_date": "05/31/2020 06:59:00",
          "content": "<p>They haven't released the source code yet, just the pre print paper a few days back that I shared above. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868845,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/31/2020 14:24:51",
          "content": "<p>OpenAI = ClosedAI</p>\n\n<p>They will likely release nothing for a while because they think \"it may be too dangerous\" lol. </p>\n\n<p>And Given the size of the model , forget about using it here ( Unless there are smaller versions)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868949,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "05/31/2020 15:38:34",
          "content": "<p>So true hahaha.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 868298,
      "author_name": "drpatrickchan",
      "author_url": "",
      "post_date": "05/31/2020 06:00:32",
      "content": "<p>Gigantic models getting much bigger. We need to think of novel approach to finetune these models on less resources.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 868379,
      "author_name": "nxrprime",
      "author_url": "",
      "post_date": "05/31/2020 07:05:29",
      "content": "<p>I said this before and I will 100 percent say it now.</p>\n\n<blockquote>\n  <p>Before it was about an elegant approach and well-written code. Now it is basically \"MY MODEL TRAINED FOR 1 WEEK ON 16 GPUS WITH 500B PARAMETERS. BEAT THAT GOOGLE!\"</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 868872,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/31/2020 14:45:42",
          "content": "<p>Language modelling is difficult, particularly predicting new task. </p>\n\n<p>So , they think the only way to overcome this, is to feed the model with as much data as possible  (That's not what I'd call \"Intelligence\" :p) </p>\n\n<p>Anyway Google has started a smart move with the Electra Model. </p>\n\n<p>But as you say, there should be room for elegant NLP papers instead of just focusing on these SOTA performance. </p>\n\n<p>Attention is a cutting Edge feature compared to the really slow and unparallelizable LSTM and both the Transformer and Bert papers were big game changers. But may be it's time to move into \"smarter\" directions than these \" Bigger and Bigger\" models that no one or very few will use in real life Applications. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868930,
          "author_name": "nxrprime",
          "author_url": "",
          "post_date": "05/31/2020 15:16:42",
          "content": "<blockquote>\n  <p>Attention is a cutting Edge feature compared to the really slow and unparallelizable LSTM and both the Transformers and Bert were big game changers. But may be it's time to move into \"smarter\" directions than these \" Bigger and Bigger\" models that no one or very few will use in real life Applications.</p>\n</blockquote>\n\n<p>I feel that we should try to look at the next big breakthrough in NLP/CV in general rather than wasting our time on these titanic models. </p>\n\n<p>I would like to see something novel like Attention crop up, or maybe the Transformer.... anything that is \"novel\". Not \"bigger\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 869065,
      "author_name": "dronych",
      "author_url": "",
      "post_date": "05/31/2020 17:53:58",
      "content": "<p>175B params - thats crazy. i will leave it here just in case someone wants to follow that <a href=\"https://github.com/openai/gpt-3\">https://github.com/openai/gpt-3</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 877743,
      "author_name": "soerendip",
      "author_url": "",
      "post_date": "06/07/2020 23:08:21",
      "content": "<p><a href=\"https://gateway.on24.com/wcc/eh/2336732/lp/2357495/burning-the-world-with-machine-learning%3A-understanding-the-carbon-footprint-of-deep-learning/\">https://gateway.on24.com/wcc/eh/2336732/lp/2357495/burning-the-world-with-machine-learning%3A-understanding-the-carbon-footprint-of-deep-learning/</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 883670,
      "author_name": "ipythonx",
      "author_url": "",
      "post_date": "06/12/2020 20:18:07",
      "content": "<p>I think people who're designing such a gigantic network should be constrained Kaggle like computing resources, like 30 hours GPU and TPU quote for their research. So, they have to think to design reasonalbe size awesome network 😅 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "867543": "Just [released](https://arxiv.org/abs/2005.14165)  .. the task agnostic, zero / one shot sounds great, the only problem is 175B parameters (~ 350 G GPU memory with 16 bit?) 😄 😄",
    "867571": "Wow! so XLM-R is just too small! \nIt looks like, that the secret of humans learning from just a few samples is just because of our large brains? ;)\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F451025%2F6bfb4878af94fc632eb0a98a19e3e9fe%2Fgpt3.png?generation=1590841027905156&amp;alt=media)",
    "867989": "GPT3 will be like :D \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F0e61c7c70114249f40754b09629b3d24%2Fimages%20-%202020-05-31T012413.281.jpeg?generation=1590868527705882&amp;alt=media)",
    "868217": "Any tutorial?",
    "868298": "Gigantic models getting much bigger. We need to think of novel approach to finetune these models on less resources.",
    "868370": "They haven't released the source code yet, just the pre print paper a few days back that I shared above.",
    "868379": "I said this before and I will 100 percent say it now.\n\n&gt; Before it was about an elegant approach and well-written code. Now it is basically \"MY MODEL TRAINED FOR 1 WEEK ON 16 GPUS WITH 500B PARAMETERS. BEAT THAT GOOGLE!\"",
    "868845": "OpenAI = ClosedAI\n\nThey will likely release nothing for a while because they think \"it may be too dangerous\" lol. \n\nAnd Given the size of the model , forget about using it here ( Unless there are smaller versions)",
    "868872": "Language modelling is difficult, particularly predicting new task. \n\nSo , they think the only way to overcome this, is to feed the model with as much data as possible  (That's not what I'd call \"Intelligence\" :p) \n\nAnyway Google has started a smart move with the Electra Model. \n\nBut as you say, there should be room for elegant NLP papers instead of just focusing on these SOTA performance. \n\nAttention is a cutting Edge feature compared to the really slow and unparallelizable LSTM and both the Transformer and Bert papers were big game changers. But may be it's time to move into \"smarter\" directions than these \" Bigger and Bigger\" models that no one or very few will use in real life Applications.",
    "868930": "&gt; Attention is a cutting Edge feature compared to the really slow and unparallelizable LSTM and both the Transformers and Bert were big game changers. But may be it's time to move into \"smarter\" directions than these \" Bigger and Bigger\" models that no one or very few will use in real life Applications.\n\nI feel that we should try to look at the next big breakthrough in NLP/CV in general rather than wasting our time on these titanic models. \n\nI would like to see something novel like Attention crop up, or maybe the Transformer.... anything that is \"novel\". Not \"bigger\".",
    "868949": "So true hahaha.",
    "869065": "175B params - thats crazy. i will leave it here just in case someone wants to follow that [https://github.com/openai/gpt-3](https://github.com/openai/gpt-3)",
    "872992": "I think that's really true. I mean even if we assume a single weight to be a single neuron we're still not nearly as big as the brain. Maybe the answer is infact bigger networks.",
    "877743": "https://gateway.on24.com/wcc/eh/2336732/lp/2357495/burning-the-world-with-machine-learning%3A-understanding-the-carbon-footprint-of-deep-learning/",
    "883670": "I think people who're designing such a gigantic network should be constrained Kaggle like computing resources, like 30 hours GPU and TPU quote for their research. So, they have to think to design reasonalbe size awesome network 😅",
    "992728": "With proper API keys, could one access APIs straight from the Kaggle notebooks?",
    "2204501": "So i made an api wrapper for the Da Vinci model provided by open-ai you can easily run it and use it in web apps and android apps. Please upvote if you like it :D."
  },
  "source": "meta"
}