{
  "id": 70715,
  "title": "Welcome!",
  "url": "/competitions/quora-insincere-questions-classification/discussion/70715",
  "author_name": "inversion",
  "post_date": "2018-11-06T18:00:29.746000",
  "votes": 37,
  "comment_count": 117,
  "views": 0,
  "content": "<p>Welcome to the Quora Insincere Questions Classification Challenge!</p>\n\n<p>In this challenge, you are tasked with developing models that identify insincere questions submitted to Quora.</p>\n\n<p>This is a Kernels-Only competition - you should familiarize yourself with the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification#Kernels-FAQ\">Kernels FAQ</a> for important information about Kernels runtime and environment constraints to be able to make a submission.</p>\n\n<p>You'll also want to read through the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/data\">Data</a> page to understand the target variable, supplementary data that's provided, etc.</p>\n\n<p>Please feel free to ask your questions in this thread.</p>\n\n<p>Good luck!</p>",
  "messages": [
    {
      "id": 416485,
      "postDate": "2018-11-06T18:00:29.747Z",
      "content": "<p>Welcome to the Quora Insincere Questions Classification Challenge!</p>\n\n<p>In this challenge, you are tasked with developing models that identify insincere questions submitted to Quora.</p>\n\n<p>This is a Kernels-Only competition - you should familiarize yourself with the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification#Kernels-FAQ\">Kernels FAQ</a> for important information about Kernels runtime and environment constraints to be able to make a submission.</p>\n\n<p>You'll also want to read through the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/data\">Data</a> page to understand the target variable, supplementary data that's provided, etc.</p>\n\n<p>Please feel free to ask your questions in this thread.</p>\n\n<p>Good luck!</p>",
      "rawMarkdown": "Welcome to the Quora Insincere Questions Classification Challenge!\n\nIn this challenge, you are tasked with developing models that identify insincere questions submitted to Quora.\n\nThis is a Kernels-Only competition - you should familiarize yourself with the [Kernels FAQ](https://www.kaggle.com/c/quora-insincere-questions-classification#Kernels-FAQ) for important information about Kernels runtime and environment constraints to be able to make a submission.\n\nYou'll also want to read through the [Data](https://www.kaggle.com/c/quora-insincere-questions-classification/data) page to understand the target variable, supplementary data that's provided, etc.\n\nPlease feel free to ask your questions in this thread.\n\nGood luck!",
      "votes": 35
    },
    {
      "id": 417390,
      "postDate": "2018-11-08T07:20:13.583Z",
      "content": "<p>Looking at the training data, am I the only one who sees lots of entries which seem to have been tagged w/o careful review?  Examples - </p>\n\n<p>0098e5b8335a62791abf,Is Dr. Steven Greer an extremely good fraudster with his claims of ufo' s and aliens?,1\n009afb0a3716d819bf7e,What did Trump and Sessions possibly gain from firing McCabe with less than 24 hours until retiring?,1\nWhy do progressive activists feel that tearing down antique historic but politically incorrect statues accomplishes a positive good?,1\ndc002c8cff10ddc797dc,Does Trump's inaction on his own justice department's indictment of Russian cyber attacks mean he has violated his oath of office?,1\ndc16793a4d38a9a27803,How do black people celebrate Martin Luther King Jr's birthday?,1\n00b92227d01a117d9c85,\"Do white people in America think they are privileged? Is that a real thing going on, or just a bad meme?\",1\ndbd98d3003044f51791a,Do Democrats have a plan to curb violence in Chicago?,1\n00c6da623540ba7d0a9d,\"What would have to be proven to show that Cohen is guilty of influence peddling, a crime? Any more than is already a matter of public record?\",1\ndc44f81813106a237458,Why doesn't China apologize for invading Tibet? They did it even before the Sino-Indian War.,1\n00c7fbcb0495b64ed7c2,\"If Donald Trump pardons himself, would that count as an admission of guilt and make it easier for the state attorney general to prosecute on related but not identical charges?\",1\n00cc8173e3c8a2a5c3fd,Can asperger's recognise sarcasm in relationships?,1 \n00df0e17c0de8d326227Do some atheists feel guilty about indoctrinating their children with atheism and not giving them hope and a moral compass?,1\ndcb2c4b6be697b16c9e3,How do I avoid losing the argument when I try to defend ethnic groups with statistically proven higher criminality rates? Am I wrong to defend these ethnic groups in the face of hard facts?,1\ndcbaf4d21e0b29592e8c,Does the conservatives’ view on climate change contradict their core conservative values?,1\ndc972c2e4f084dc8fdb8,\"Why do so many Americans get puzzled when I ask for a serviette in their restaurants, but know right away what I am talking about when I ask for a Napkin (aka diaper)?\",1\n00df34daff54eef050f7What is it like to be a gay in IIT Madras?,1</p>\n\n<p>But here we have  : dbae47ac9a7adf7f178cWhat is it like to be gay in Southern California?,0</p>\n\n<p>(So - gay as an adjective is OK, but not as a noun?)</p>\n\n<p>This is about 10 % of the 1's, near as I can reckon.    Yes, these are ham-fisted questions about delicate topics.   Yes, they could be worded more gently.   But they do not harass nor disparage, they certainly appear to be in good faith and seeking real answers.</p>\n\n<p>There's a real garbage-in garbage-out problem here.   The dataset appears not to have been human-curated at all, but rather from an algorithm which is surfacing false positives.   Or from a curator who is all nerve endings about certain issues.</p>\n\n<p>I was very excited to enter (and win!!!) this competition but I am stopped dead in my tracks by bad training data.  My complaint isn't about politics, I could not describe to a human with a PhD in English Literature how to arrive at the conclusions given in this sample.</p>\n\n<p>Trying to code it is a fool's errand, to put it mildly.</p>\n\n<p>Not convinced?   How about this one - </p>\n\n<p>dcbfe10facdd66c91eca,Why aren't there any women in the Navy SEALS?,1</p>\n\n<p>Explain that one.   To the US Navy.    <a href=\"https://www.navytimes.com/news/your-navy/2018/02/16/two-women-could-enter-navy-special-operations-training-this-year/\">https://www.navytimes.com/news/your-navy/2018/02/16/two-women-could-enter-navy-special-operations-training-this-year/</a></p>",
      "rawMarkdown": "Looking at the training data, am I the only one who sees lots of entries which seem to have been tagged w/o careful review?  Examples - \n\n0098e5b8335a62791abf,Is Dr. Steven Greer an extremely good fraudster with his claims of ufo' s and aliens?,1\n009afb0a3716d819bf7e,What did Trump and Sessions possibly gain from firing McCabe with less than 24 hours until retiring?,1\nWhy do progressive activists feel that tearing down antique historic but politically incorrect statues accomplishes a positive good?,1\ndc002c8cff10ddc797dc,Does Trump's inaction on his own justice department's indictment of Russian cyber attacks mean he has violated his oath of office?,1\ndc16793a4d38a9a27803,How do black people celebrate Martin Luther King Jr's birthday?,1\n00b92227d01a117d9c85,\"Do white people in America think they are privileged? Is that a real thing going on, or just a bad meme?\",1\ndbd98d3003044f51791a,Do Democrats have a plan to curb violence in Chicago?,1\n00c6da623540ba7d0a9d,\"What would have to be proven to show that Cohen is guilty of influence peddling, a crime? Any more than is already a matter of public record?\",1\ndc44f81813106a237458,Why doesn't China apologize for invading Tibet? They did it even before the Sino-Indian War.,1\n00c7fbcb0495b64ed7c2,\"If Donald Trump pardons himself, would that count as an admission of guilt and make it easier for the state attorney general to prosecute on related but not identical charges?\",1\n00cc8173e3c8a2a5c3fd,Can asperger's recognise sarcasm in relationships?,1 \n00df0e17c0de8d326227Do some atheists feel guilty about indoctrinating their children with atheism and not giving them hope and a moral compass?,1\ndcb2c4b6be697b16c9e3,How do I avoid losing the argument when I try to defend ethnic groups with statistically proven higher criminality rates? Am I wrong to defend these ethnic groups in the face of hard facts?,1\ndcbaf4d21e0b29592e8c,Does the conservatives’ view on climate change contradict their core conservative values?,1\ndc972c2e4f084dc8fdb8,\"Why do so many Americans get puzzled when I ask for a serviette in their restaurants, but know right away what I am talking about when I ask for a Napkin (aka diaper)?\",1\n00df34daff54eef050f7What is it like to be a gay in IIT Madras?,1\n\n\nBut here we have  : dbae47ac9a7adf7f178cWhat is it like to be gay in Southern California?,0\n\n(So - gay as an adjective is OK, but not as a noun?)\n\nThis is about 10 % of the 1's, near as I can reckon.    Yes, these are ham-fisted questions about delicate topics.   Yes, they could be worded more gently.   But they do not harass nor disparage, they certainly appear to be in good faith and seeking real answers.\n\nThere's a real garbage-in garbage-out problem here.   The dataset appears not to have been human-curated at all, but rather from an algorithm which is surfacing false positives.   Or from a curator who is all nerve endings about certain issues.\n\nI was very excited to enter (and win!!!) this competition but I am stopped dead in my tracks by bad training data.  My complaint isn't about politics, I could not describe to a human with a PhD in English Literature how to arrive at the conclusions given in this sample.\n\nTrying to code it is a fool's errand, to put it mildly.\n\nNot convinced?   How about this one - \n\ndcbfe10facdd66c91eca,Why aren't there any women in the Navy SEALS?,1\n\nExplain that one.   To the US Navy.    https://www.navytimes.com/news/your-navy/2018/02/16/two-women-could-enter-navy-special-operations-training-this-year/\n\n\n\n\n\n\n",
      "votes": 22,
      "replies": [
        {
          "id": 417641,
          "postDate": "2018-11-08T15:07:04.427Z",
          "content": "<p>I feel like some of the questions you highlighted are biased or too badly-worded, thus flagged as insincere.</p>\n\n<p>Yet I have to recognize that some questions should not be highlighted as insincere (a gay vs. gay is the best example).</p>",
          "rawMarkdown": "I feel like some of the questions you highlighted are biased or too badly-worded, thus flagged as insincere.\n\nYet I have to recognize that some questions should not be highlighted as insincere (a gay vs. gay is the best example).",
          "votes": 1
        },
        {
          "id": 417804,
          "postDate": "2018-11-08T19:53:43.393Z",
          "content": "<p>I had the same concerns after looking through some of the questions that were flagged/not flagged. You focus on things that seem incorrectly flagged, but there are also a lot of things they've obviously missed flagging like single word/nonsense kinds of things. I'm concerned we're not going to be able to train anything better than what they've already got when the training data is so poorly curated. I was also wondering whether the data set we're ultimately tested on will be of similar quality.</p>",
          "rawMarkdown": "I had the same concerns after looking through some of the questions that were flagged/not flagged. You focus on things that seem incorrectly flagged, but there are also a lot of things they've obviously missed flagging like single word/nonsense kinds of things. I'm concerned we're not going to be able to train anything better than what they've already got when the training data is so poorly curated. I was also wondering whether the data set we're ultimately tested on will be of similar quality.",
          "votes": 5
        },
        {
          "id": 417819,
          "postDate": "2018-11-08T20:24:42.973Z",
          "content": "<p>I hadn't even looked at the false negatives.   I suppose if the final data testing data is of low quality, that should lower everyone's score by a similar amount, so the final ranking may still be accurate.  But maybe not.</p>\n\n<p>Are the false negatives all simply malformed/meaningless?  I don't see that the target is defined to include those (<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/data\">https://www.kaggle.com/c/quora-insincere-questions-classification/data</a>)</p>\n\n<p>But too much noise in the training data - yea, can't really progress much with garbage-in.</p>\n\n<p>Should we all gang up and try to clean up the data?   There are 100K flagged records, it takes about 2 seconds to check each, if 50 people spend an hour each, and then another 1/2 hour to check each other's work - that would do it.</p>",
          "rawMarkdown": "I hadn't even looked at the false negatives.   I suppose if the final data testing data is of low quality, that should lower everyone's score by a similar amount, so the final ranking may still be accurate.  But maybe not.\n\nAre the false negatives all simply malformed/meaningless?  I don't see that the target is defined to include those (https://www.kaggle.com/c/quora-insincere-questions-classification/data)\n\nBut too much noise in the training data - yea, can't really progress much with garbage-in.\n\nShould we all gang up and try to clean up the data?   There are 100K flagged records, it takes about 2 seconds to check each, if 50 people spend an hour each, and then another 1/2 hour to check each other's work - that would do it.\n",
          "votes": 2
        },
        {
          "id": 417902,
          "postDate": "2018-11-09T00:38:37.040Z",
          "content": "<p>Right, it's still fair if we've all got the same data, but probably less fun and useful.</p>\n\n<p>And you're right that the description didn't specifically include malformed/meaningless things in the insincere category. But if you look at the really short questions, for example, it looks like they are actually trying to flag malformed questions:</p>\n\n<p>sincere: 'Is God 42?', 'In Islam?', 'I 12?', 'How can I?', 'Am I fake?', 'Wat is 1A?', 'What sexy?', 'Why is 16?', 'What meow?', 'Why is αθ?', 'Hello sir?', 'IS 1+1 21?', 'ESR is 35?'</p>\n\n<p>insincere: 'Quora: ', 'Is ', 'Dowry: ', 'If ', 'To Quora:', 'Sexism: ', 'Hungary: ', 'History: ', 'Islam: ', 'Politics: ', 'Feminism: ', '\"', 'Germany: ', 'Nepal: ', 'India: ', \"I'm an \", 'Wicca: ', 'Bye Bye?', 'Incest: ', 'Do ', 'Poland: ', 'W', 'Why is ', 'Identity: '</p>\n\n<p>Anyway, I don't really have a sense at this point of whether the false positives or false negatives are a bigger issue overall.  </p>",
          "rawMarkdown": "Right, it's still fair if we've all got the same data, but probably less fun and useful.\n\nAnd you're right that the description didn't specifically include malformed/meaningless things in the insincere category. But if you look at the really short questions, for example, it looks like they are actually trying to flag malformed questions:\n\nsincere: 'Is God 42?', 'In Islam?', 'I 12?', 'How can I?', 'Am I fake?', 'Wat is 1A?', 'What sexy?', 'Why is 16?', 'What meow?', 'Why is αθ?', 'Hello sir?', 'IS 1+1 21?', 'ESR is 35?'\n \ninsincere: 'Quora: ', 'Is ', 'Dowry: ', 'If ', 'To Quora:', 'Sexism: ', 'Hungary: ', 'History: ', 'Islam: ', 'Politics: ', 'Feminism: ', '\"', 'Germany: ', 'Nepal: ', 'India: ', \"I'm an \", 'Wicca: ', 'Bye Bye?', 'Incest: ', 'Do ', 'Poland: ', 'W', 'Why is ', 'Identity: '\n\nAnyway, I don't really have a sense at this point of whether the false positives or false negatives are a bigger issue overall.  ",
          "votes": 3
        },
        {
          "id": 420230,
          "postDate": "2018-11-13T10:14:16.467Z",
          "content": "<p>I have got a few examples which I believe should be flagged, but i guess those flagged are sufficient to be used for training now.</p>\n\n<p>0015a0bfd83e1fa30b73 - Which black superhero can beat Superman or is as strong as him?\n00c25e8c423fa71efcb2 - Why does this girl of a different race than me think I am racist just becausedo not find her physically attractive?\n0089e5eb91ca56c62a8b - Should I be concerned if I saw a white person with a black child?</p>\n\n<p>Pleas share your thoughts from <a href=\"/inversion\">@inversion</a></p>",
          "rawMarkdown": "I have got a few examples which I believe should be flagged, but i guess those flagged are sufficient to be used for training now.\n\n0015a0bfd83e1fa30b73 - Which black superhero can beat Superman or is as strong as him?\n00c25e8c423fa71efcb2 - Why does this girl of a different race than me think I am racist just becausedo not find her physically attractive?\n0089e5eb91ca56c62a8b - Should I be concerned if I saw a white person with a black child?\n\nPleas share your thoughts from @inversion",
          "votes": 1
        },
        {
          "id": 430132,
          "postDate": "2018-11-29T22:02:46.150Z",
          "content": "<p>@IridiumBlue: So - gay as an adjective is OK, but not as a noun?</p>\n\n<p>That seems perfectly reasonable to me. In my experience, in English, \"gay\" is normally an adjective (though it may occasionally be used as a noun), but in Trollspeak, it is commonly used as a noun (and maybe not often as an adjective, except in phrases like \"that's so gay\").</p>",
          "rawMarkdown": "@IridiumBlue: So - gay as an adjective is OK, but not as a noun?\n\nThat seems perfectly reasonable to me. In my experience, in English, \"gay\" is normally an adjective (though it may occasionally be used as a noun), but in Trollspeak, it is commonly used as a noun (and maybe not often as an adjective, except in phrases like \"that's so gay\").",
          "votes": 1
        },
        {
          "id": 430324,
          "postDate": "2018-11-30T07:56:11.260Z",
          "content": "<p>This is indeed a serious problem, quite a lot of the questions are false positive/ false negatives, which makes it quite difficult to train an efficient algorithm. We need to have a better dataset, otherwise our algorithms will not be so useful, and only be approximating the way the Quora team( combining  algorithms and manual labeling) label the insincere questions, not so practical then.\nI wish they can make an effort trying to clean some data.</p>",
          "rawMarkdown": "This is indeed a serious problem, quite a lot of the questions are false positive/ false negatives, which makes it quite difficult to train an efficient algorithm. We need to have a better dataset, otherwise our algorithms will not be so useful, and only be approximating the way the Quora team( combining  algorithms and manual labeling) label the insincere questions, not so practical then.\nI wish they can make an effort trying to clean some data.",
          "votes": -1
        }
      ]
    },
    {
      "id": 422736,
      "postDate": "2018-11-16T17:42:23.147Z",
      "content": "<p>Seems most of us use CUDNN package which has strong randomness for each run (but very fast). Is that possible in stage 2, you run the model a few times and take an average?</p>",
      "rawMarkdown": "Seems most of us use CUDNN package which has strong randomness for each run (but very fast). Is that possible in stage 2, you run the model a few times and take an average?",
      "votes": 7,
      "replies": [
        {
          "id": 439077,
          "postDate": "2018-12-14T17:34:39.127Z",
          "content": "<p>Please answer this <a href=\"/inversion\">@inversion</a> , <a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>We are experiencing huge randomness and LB is not very reliable </p>",
          "rawMarkdown": "Please answer this @inversion , @juliaelliott \n\nWe are experiencing huge randomness and LB is not very reliable "
        },
        {
          "id": 439106,
          "postDate": "2018-12-14T18:40:27.987Z",
          "content": "<p>It would be unfair for those though who try to tune their model to be as stable as possible.</p>",
          "rawMarkdown": "It would be unfair for those though who try to tune their model to be as stable as possible.",
          "votes": 3
        },
        {
          "id": 440595,
          "postDate": "2018-12-17T18:40:21.487Z",
          "content": "<p>Kernels will only be run once.</p>",
          "rawMarkdown": "Kernels will only be run once.",
          "votes": 3
        },
        {
          "id": 441637,
          "postDate": "2018-12-18T21:50:38.533Z",
          "content": "<p>Thanks <a href=\"/inversion\">@inversion</a>, makes total sense to me. One more question though: As the test size for private LB will be larger, the kernel will need more time to run. Do we need to make an estimate for that, or is it enough if the current public LB version runs within 2 hours?</p>",
          "rawMarkdown": "Thanks @inversion, makes total sense to me. One more question though: As the test size for private LB will be larger, the kernel will need more time to run. Do we need to make an estimate for that, or is it enough if the current public LB version runs within 2 hours?"
        },
        {
          "id": 441672,
          "postDate": "2018-12-18T23:09:30.007Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Kernels will need to take into consideration the additional processing and inference time of the larger Test dataset. That will need to fit into the 2 / 6 hour constraints.</p>",
          "rawMarkdown": "@philippsinger Kernels will need to take into consideration the additional processing and inference time of the larger Test dataset. That will need to fit into the 2 / 6 hour constraints."
        }
      ]
    },
    {
      "id": 416988,
      "postDate": "2018-11-07T15:13:24.570Z",
      "content": "<p>Great to have another NLP-related competition on Kaggle! Surely going to give this a go!</p>",
      "rawMarkdown": "Great to have another NLP-related competition on Kaggle! Surely going to give this a go!",
      "votes": 7
    },
    {
      "id": 416747,
      "postDate": "2018-11-07T08:04:31.293Z",
      "content": "<p>Can the docker image change during the course of the competition? E.g. can it happen that some interesting library will not be available at the beginning, but and made available later? If so, do I have to monitor commits to the official docker image to make myself aware of this?</p>",
      "rawMarkdown": "Can the docker image change during the course of the competition? E.g. can it happen that some interesting library will not be available at the beginning, but and made available later? If so, do I have to monitor commits to the official docker image to make myself aware of this?\n",
      "votes": 8,
      "replies": [
        {
          "id": 417904,
          "postDate": "2018-11-09T00:53:22.167Z",
          "content": "<p>Great question! We expect the Kaggle Docker image to have updates over the course of the competition, and there are no restrictions to using the latest updates.</p>",
          "rawMarkdown": "Great question! We expect the Kaggle Docker image to have updates over the course of the competition, and there are no restrictions to using the latest updates.",
          "votes": 3
        },
        {
          "id": 417956,
          "postDate": "2018-11-09T03:49:47.550Z",
          "content": "<p>How long does it usually take for a pull request to the docker-python repository to be added to the kernel's image?</p>",
          "rawMarkdown": "How long does it usually take for a pull request to the docker-python repository to be added to the kernel's image?"
        }
      ]
    },
    {
      "id": 416527,
      "postDate": "2018-11-06T20:14:01.027Z",
      "content": "<p>In stage 2 of the competition you rerun the kernels on a test set of approx 376k vs approx 56k in the first stage. Is there any adjustment made to the available run times in the kernels, or do they still have to run in 6 hours (or 2 hours for GPU)?</p>",
      "rawMarkdown": "In stage 2 of the competition you rerun the kernels on a test set of approx 376k vs approx 56k in the first stage. Is there any adjustment made to the available run times in the kernels, or do they still have to run in 6 hours (or 2 hours for GPU)?",
      "votes": 5,
      "replies": [
        {
          "id": 416694,
          "postDate": "2018-11-07T05:34:50.440Z",
          "content": "<p>The run times (as you’ve stated) will remain the same.</p>",
          "rawMarkdown": "The run times (as you’ve stated) will remain the same.",
          "votes": 4
        },
        {
          "id": 417190,
          "postDate": "2018-11-07T23:11:47.253Z",
          "content": "<p>Thanks for confirming.</p>",
          "rawMarkdown": "Thanks for confirming.",
          "votes": 1
        },
        {
          "id": 437209,
          "postDate": "2018-12-11T14:49:45.260Z",
          "content": "<p><code>\nyour runtime of ... minutes exceeds the CPU kernel max of 120 minutes\n</code>\nwhat exactly happened here?</p>",
          "rawMarkdown": "```\nyour runtime of ... minutes exceeds the CPU kernel max of 120 minutes\n```\nwhat exactly happened here?"
        }
      ]
    },
    {
      "id": 424161,
      "postDate": "2018-11-19T17:04:03.737Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> Are vectors and other models from spaCy fair game??</p>\n\n<p>Other people have asked same thing here: <a href=\"https://www.kaggle.com/jpmiller/bonus-vectors-with-spacy/comments\">https://www.kaggle.com/jpmiller/bonus-vectors-with-spacy/comments</a></p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "@inversion Are vectors and other models from spaCy fair game??\n\nOther people have asked same thing here: https://www.kaggle.com/jpmiller/bonus-vectors-with-spacy/comments\n\nThanks!",
      "votes": 6,
      "replies": [
        {
          "id": 424347,
          "postDate": "2018-11-20T00:54:50.357Z",
          "content": "<p><a href=\"/tezdhar\">@tezdhar</a> I'd extend this question even more: Are vectors and other models from available Docker image fair game?</p>\n\n<p><a href=\"/inversion\">@inversion</a> sorry to ping you again, but please answer on these questions asap. It just might save a lot of time for all participants.</p>",
          "rawMarkdown": "@tezdhar I'd extend this question even more: Are vectors and other models from available Docker image fair game?\n\n@inversion sorry to ping you again, but please answer on these questions asap. It just might save a lot of time for all participants.",
          "votes": 2
        },
        {
          "id": 426039,
          "postDate": "2018-11-22T13:45:28.560Z",
          "content": "<p><a href=\"/inversion\">@inversion</a> Sorry to also ping you again. But a reply to that question is pretty crucial.</p>",
          "rawMarkdown": "@inversion Sorry to also ping you again. But a reply to that question is pretty crucial.",
          "votes": 2
        },
        {
          "id": 426119,
          "postDate": "2018-11-22T16:48:50.820Z",
          "content": "<p>Tools that are available in the Docker image are fair game.</p>",
          "rawMarkdown": "Tools that are available in the Docker image are fair game.",
          "votes": 8
        },
        {
          "id": 443744,
          "postDate": "2018-12-22T10:11:28.020Z",
          "content": "<p>Now that the current docker has fastai v1 which includes the WT103 (wiki text) language model (not among white listed), is this answer still valid?</p>",
          "rawMarkdown": "Now that the current docker has fastai v1 which includes the WT103 (wiki text) language model (not among white listed), is this answer still valid?"
        },
        {
          "id": 443768,
          "postDate": "2018-12-22T11:36:06.607Z",
          "content": "<p>Hi <a href=\"/abedkhooli\">@abedkhooli</a>,\nAre you sure that it's possible to use fasta1 pre-trained LM without internet access? Shouldn't it be downloaded separately from the package?</p>",
          "rawMarkdown": "Hi @abedkhooli,\nAre you sure that it's possible to use fasta1 pre-trained LM without internet access? Shouldn't it be downloaded separately from the package?"
        },
        {
          "id": 443800,
          "postDate": "2018-12-22T13:10:57.763Z",
          "content": "<p>Technically speaking, the model gets downloaded (at run time) if referenced in the kernel code but need not add as part of the datasets, that's why I asked for a definite answer from <a href=\"/inversion\">@inversion</a> </p>",
          "rawMarkdown": "Technically speaking, the model gets downloaded (at run time) if referenced in the kernel code but need not add as part of the datasets, that's why I asked for a definite answer from @inversion "
        },
        {
          "id": 443805,
          "postDate": "2018-12-22T13:33:50.283Z",
          "content": "<p>You cannot access Internet through runtime in scoring afaik.</p>",
          "rawMarkdown": "You cannot access Internet through runtime in scoring afaik.",
          "votes": 1
        },
        {
          "id": 445599,
          "postDate": "2018-12-26T19:39:54.050Z",
          "content": "<p>It runs but after commit the submission link is disabled. The tooltip reads:\n\"You cannot use internet access for this competition. Your runtime of 134 minutes exceeds the GPU kernel max of 120 minutes.\"\nSo, apparently there is limit on tools and compute.</p>",
          "rawMarkdown": "It runs but after commit the submission link is disabled. The tooltip reads:\n\"You cannot use internet access for this competition. Your runtime of 134 minutes exceeds the GPU kernel max of 120 minutes.\"\nSo, apparently there is limit on tools and compute.\n"
        }
      ]
    },
    {
      "id": 417221,
      "postDate": "2018-11-08T00:44:49.207Z",
      "content": "<p>Hi inversion, can we have more digits shown on the LB? Say 3 or 4 like most other competitions. Thanks</p>",
      "rawMarkdown": "Hi inversion, can we have more digits shown on the LB? Say 3 or 4 like most other competitions. Thanks",
      "votes": 5,
      "replies": [
        {
          "id": 417912,
          "postDate": "2018-11-09T01:07:09.013Z",
          "content": "<p>Things saturated rather quickly! I increased digits to 3, and will keep it there for most of the competition. </p>\n\n<p>This is one where I hope people rely heavily on local validation anyway. :-)</p>",
          "rawMarkdown": "Things saturated rather quickly! I increased digits to 3, and will keep it there for most of the competition. \n\nThis is one where I hope people rely heavily on local validation anyway. :-)",
          "votes": 6
        }
      ]
    },
    {
      "id": 443853,
      "postDate": "2018-12-22T15:30:49.510Z",
      "content": "<p>my first kaggle competition! go go go</p>",
      "rawMarkdown": "my first kaggle competition! go go go",
      "votes": 3
    },
    {
      "id": 418186,
      "postDate": "2018-11-09T12:51:33.093Z",
      "content": "<p>Hi Inversion,</p>\n\n<p>External Data is not allowed, fine. \nBut can you please add Wiki-text/allow us to use some standard datasets from which we can train our own language models and do transfer learning.\nWhat is the point behind not allowing language models? Isn't it a loss for the organizer when participants are handicapped by not being allowed to use state of art capability?\nNot to mention amount of learning that will be missed due to this?</p>",
      "rawMarkdown": "Hi Inversion,\n\nExternal Data is not allowed, fine. \nBut can you please add Wiki-text/allow us to use some standard datasets from which we can train our own language models and do transfer learning.\nWhat is the point behind not allowing language models? Isn't it a loss for the organizer when participants are handicapped by not being allowed to use state of art capability?\nNot to mention amount of learning that will be missed due to this?",
      "votes": 3,
      "replies": [
        {
          "id": 418453,
          "postDate": "2018-11-09T23:04:59.293Z",
          "content": "<p>I had the same concern earlier, as for standard datasets and language models, we do have use of all the packages and associated data that are part of Kaggle Docker.   </p>\n\n<p>See <a href=\"https://github.com/Kaggle/docker-python/blob/master/Dockerfile\">https://github.com/Kaggle/docker-python/blob/master/Dockerfile</a>, especially this line : </p>\n\n<p>python -m nltk.downloader -d /usr/share/nltk_data abc alpino averaged_perceptron_tagger \\\n    basque_grammars biocreative_ppi bllip_wsj_no_aux \\\n    book_grammars brown brown_tei cess_cat cess_esp chat80 city_database cmudict \\\n    comtrans conll2000 conll2002 conll2007 crubadan dependency_treebank \\\n    europarl_raw floresta gazetteers genesis gutenberg \\\n    ieer inaugural indian jeita kimmo knbc large_grammars lin_thesaurus mac_morpho machado \\\n    masc_tagged maxent_ne_chunker maxent_treebank_pos_tagger moses_sample movie_reviews \\\n    mte_teip5 names nps_chat omw opinion_lexicon paradigms \\\n    pil pl196x porter_test ppattach problem_reports product_reviews_1 product_reviews_2 propbank \\\n    pros_cons ptb punkt qc reuters rslp rte sample_grammars semcor senseval sentence_polarity \\\n    sentiwordnet shakespeare sinica_treebank smultron snowball_data spanish_grammars \\\n    state_union stopwords subjectivity swadesh switchboard tagsets timit toolbox treebank \\\n    twitter_samples udhr2 udhr unicode_samples universal_tagset universal_treebanks_v20 \\\nvader_lexicon verbnet webtext word2vec_sample wordnet wordnet_ic words ycoe &amp;&amp; \\</p>\n\n<p>So basically any package you might want, and if you need another - I believe the contest runners are open to suggestions to add them.</p>",
          "rawMarkdown": "I had the same concern earlier, as for standard datasets and language models, we do have use of all the packages and associated data that are part of Kaggle Docker.   \n\nSee https://github.com/Kaggle/docker-python/blob/master/Dockerfile, especially this line : \n\npython -m nltk.downloader -d /usr/share/nltk_data abc alpino averaged_perceptron_tagger \\\n    basque_grammars biocreative_ppi bllip_wsj_no_aux \\\n    book_grammars brown brown_tei cess_cat cess_esp chat80 city_database cmudict \\\n    comtrans conll2000 conll2002 conll2007 crubadan dependency_treebank \\\n    europarl_raw floresta gazetteers genesis gutenberg \\\n    ieer inaugural indian jeita kimmo knbc large_grammars lin_thesaurus mac_morpho machado \\\n    masc_tagged maxent_ne_chunker maxent_treebank_pos_tagger moses_sample movie_reviews \\\n    mte_teip5 names nps_chat omw opinion_lexicon paradigms \\\n    pil pl196x porter_test ppattach problem_reports product_reviews_1 product_reviews_2 propbank \\\n    pros_cons ptb punkt qc reuters rslp rte sample_grammars semcor senseval sentence_polarity \\\n    sentiwordnet shakespeare sinica_treebank smultron snowball_data spanish_grammars \\\n    state_union stopwords subjectivity swadesh switchboard tagsets timit toolbox treebank \\\n    twitter_samples udhr2 udhr unicode_samples universal_tagset universal_treebanks_v20 \\\nvader_lexicon verbnet webtext word2vec_sample wordnet wordnet_ic words ycoe &amp;&amp; \\\n\nSo basically any package you might want, and if you need another - I believe the contest runners are open to suggestions to add them.",
          "votes": 3
        }
      ]
    },
    {
      "id": 417985,
      "postDate": "2018-11-09T04:36:46.333Z",
      "content": "<p>Hello, this challenge is forbiden for using external data, but if pretrained model or transfer learning is possible in this challenge?</p>",
      "rawMarkdown": "Hello, this challenge is forbiden for using external data, but if pretrained model or transfer learning is possible in this challenge?",
      "votes": 3,
      "replies": [
        {
          "id": 426118,
          "postDate": "2018-11-22T16:46:15.637Z",
          "content": "<p>No, since there would be no way incorporate these within the competition constraints (i.e., if you create a kernel that uses an additional dataset, you won't be able to submit the kernel output).</p>",
          "rawMarkdown": "No, since there would be no way incorporate these within the competition constraints (i.e., if you create a kernel that uses an additional dataset, you won't be able to submit the kernel output)."
        }
      ]
    },
    {
      "id": 417161,
      "postDate": "2018-11-07T21:44:50.890Z",
      "content": "<p>Hi - can we work on multiple kernels (as long as we make only one submission) ?</p>",
      "rawMarkdown": "Hi - can we work on multiple kernels (as long as we make only one submission) ?\n",
      "votes": 3,
      "replies": [
        {
          "id": 417907,
          "postDate": "2018-11-09T00:59:25.333Z",
          "content": "<p>Hi @Sundaresh -</p>\n\n<p>Yes, you can create multiple Kernels, but at the end of the competition you will only be able to select two for final scoring.</p>",
          "rawMarkdown": "Hi @Sundaresh -\n\nYes, you can create multiple Kernels, but at the end of the competition you will only be able to select two for final scoring.",
          "votes": 1
        }
      ]
    },
    {
      "id": 417173,
      "postDate": "2018-11-07T22:11:32.567Z",
      "content": "<p>What are the constrains on local data?  I get that we are limited in our access to large external datasets, but suppose we have an array of 100 magic numbers that makes our kernel a winner?  10,000?   1,000,000 ?   </p>",
      "rawMarkdown": "What are the constrains on local data?  I get that we are limited in our access to large external datasets, but suppose we have an array of 100 magic numbers that makes our kernel a winner?  10,000?   1,000,000 ?   ",
      "votes": 4,
      "replies": [
        {
          "id": 417910,
          "postDate": "2018-11-09T01:02:53.720Z",
          "content": "<p>You can utilize any data that is generated during the Kernel run. Most people, for example, will generate new data as part of the feature creation process. </p>",
          "rawMarkdown": "You can utilize any data that is generated during the Kernel run. Most people, for example, will generate new data as part of the feature creation process. "
        },
        {
          "id": 423712,
          "postDate": "2018-11-18T22:59:56.467Z",
          "content": "<p>Code is not generated during a Kernel run.  I think that the OP is asking is 'what stops me using FAR more compute offline to pre-train a network [with arbitrary external data] and write a script that extracts all its parameters as constants in code, which I then paste into the kernel and instantiate my model weights from?</p>\n\n<p>Obviously this is intended to be disallowed, but its somewhat subjective where the boundary is.  For instance model parameters vs model hyper-parameters - I'm pretty sure most high ranked kernels will be using hyper-parameter values that result from far more extensive out-of-kernel training.</p>\n\n<p>In general it's not possible to police this sensibly, but if the rules were clear on the appropriate demarkation, it could at least make the final top-scoring kernels subject to a manual within-the-spirit test...</p>",
          "rawMarkdown": "Code is not generated during a Kernel run.  I think that the OP is asking is 'what stops me using FAR more compute offline to pre-train a network [with arbitrary external data] and write a script that extracts all its parameters as constants in code, which I then paste into the kernel and instantiate my model weights from?\n\nObviously this is intended to be disallowed, but its somewhat subjective where the boundary is.  For instance model parameters vs model hyper-parameters - I'm pretty sure most high ranked kernels will be using hyper-parameter values that result from far more extensive out-of-kernel training.\n\nIn general it's not possible to police this sensibly, but if the rules were clear on the appropriate demarkation, it could at least make the final top-scoring kernels subject to a manual within-the-spirit test...",
          "votes": 8
        },
        {
          "id": 434630,
          "postDate": "2018-12-06T17:42:03.347Z",
          "content": "<p>This is really tricky. eg: a list of stop words - shall one use it - would <em>easily</em> fit in memory and in code - say a list or set with 150~ or so items. Does it configure as \"unallowed use of external data\"? </p>",
          "rawMarkdown": "This is really tricky. eg: a list of stop words - shall one use it - would *easily* fit in memory and in code - say a list or set with 150~ or so items. Does it configure as \"unallowed use of external data\"? ",
          "votes": 1
        },
        {
          "id": 443466,
          "postDate": "2018-12-21T17:40:49.550Z",
          "content": "<p>I'm worried if this still hasn't been addressed properly. There are so many people using magic number lists and misspell dictionaries for preprocessing functions, which are already predefined and copy pasted into the kernel. In theory the majority of these could be written out and thought of by anyone with any idea of preprocessing, but if it falls under external data sources then kernels will need to be rewritten. Are written lists / dictionaries considered as external? If not then is there a limit to these? Would be great to know. </p>",
          "rawMarkdown": "I'm worried if this still hasn't been addressed properly. There are so many people using magic number lists and misspell dictionaries for preprocessing functions, which are already predefined and copy pasted into the kernel. In theory the majority of these could be written out and thought of by anyone with any idea of preprocessing, but if it falls under external data sources then kernels will need to be rewritten. Are written lists / dictionaries considered as external? If not then is there a limit to these? Would be great to know. ",
          "votes": 4
        },
        {
          "id": 458591,
          "postDate": "2019-01-20T02:07:00.683Z",
          "content": "<blockquote>\n  <p>There are so many people using magic number lists and misspell dictionaries for preprocessing functions, which are already predefined and copy pasted into the kernel. </p>\n</blockquote>\n\n<p>I totally agree with you. </p>\n\n<blockquote>\n  <p>Are written lists / dictionaries considered as external? If not then is there a limit to these? Would be great to know.  </p>\n</blockquote>\n\n<p>I am worried too.  </p>",
          "rawMarkdown": "&gt; There are so many people using magic number lists and misspell dictionaries for preprocessing functions, which are already predefined and copy pasted into the kernel. \n\nI totally agree with you. \n\n&gt; Are written lists / dictionaries considered as external? If not then is there a limit to these? Would be great to know.  \n\nI am worried too.  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 417143,
      "postDate": "2018-11-07T20:26:02.900Z",
      "content": "<p>can i use packages installed from pip? where i can see list of packages that i can use?</p>",
      "rawMarkdown": "can i use packages installed from pip? where i can see list of packages that i can use?",
      "votes": 4,
      "replies": [
        {
          "id": 417905,
          "postDate": "2018-11-09T00:56:41.133Z",
          "content": "<p>You will not be able to make submissions if you <code>pip install</code> libraries into your kernel.</p>\n\n<p>You can see the list of available packages in the Kaggle Docker image here:</p>\n\n<p><a href=\"https://github.com/Kaggle/docker-python\">https://github.com/Kaggle/docker-python</a></p>",
          "rawMarkdown": "You will not be able to make submissions if you `pip install` libraries into your kernel.\n\nYou can see the list of available packages in the Kaggle Docker image here:\n\nhttps://github.com/Kaggle/docker-python",
          "votes": 2
        }
      ]
    },
    {
      "id": 446977,
      "postDate": "2018-12-29T00:01:18.370Z",
      "content": "<p>It would be great to get an answer to this question:</p>\n\n<p><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75816#446311\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75816#446311</a></p>",
      "rawMarkdown": "It would be great to get an answer to this question:\n\nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75816#446311",
      "votes": 1,
      "replies": [
        {
          "id": 462417,
          "postDate": "2019-01-28T09:11:15.757Z",
          "content": "<p>The essential problem facing the Kaggle and Quora ppl is that 'code' and 'data' are artificial distinctions.   Code is data.   </p>\n\n<p>So the question reduces to : How big can our code be?  And more precisely, how big can our code be at maximum entropy (compressed to the max.)</p>\n\n<p>They don't want to answer this question, and I suppose I don't blame them.   If they declare it to be, say, 1 Gig, people will start using all that space to import 'data' (in the loose sense of the term.)</p>\n\n<p>If they make it too low, they might clip somebody who is using their own hand-rolled framework comprising 100,000 lines.</p>\n\n<p>So they opt for silence, so nobody can game the restriction.</p>\n\n<p>The signal that sends to us participants is  : Be careful.   A list of the five parts of speech, \"Noun, Adjective, etc.\" is obviously OK.    A list of 20 dirty words is probably OK.   A stop-word list is big - so watch out and use one already in the kernel.    </p>\n\n<p>A list of 1000 toxic verbs?   Thin ice.</p>\n\n<p>200 MB of matrix coefficients?   Red light.</p>",
          "rawMarkdown": "The essential problem facing the Kaggle and Quora ppl is that 'code' and 'data' are artificial distinctions.   Code is data.   \n\nSo the question reduces to : How big can our code be?  And more precisely, how big can our code be at maximum entropy (compressed to the max.)\n\nThey don't want to answer this question, and I suppose I don't blame them.   If they declare it to be, say, 1 Gig, people will start using all that space to import 'data' (in the loose sense of the term.)\n\nIf they make it too low, they might clip somebody who is using their own hand-rolled framework comprising 100,000 lines.\n\nSo they opt for silence, so nobody can game the restriction.\n\nThe signal that sends to us participants is  : Be careful.   A list of the five parts of speech, \"Noun, Adjective, etc.\" is obviously OK.    A list of 20 dirty words is probably OK.   A stop-word list is big - so watch out and use one already in the kernel.    \n\nA list of 1000 toxic verbs?   Thin ice.\n\n200 MB of matrix coefficients?   Red light.\n\n\n\n\n\n"
        }
      ]
    },
    {
      "id": 433783,
      "postDate": "2018-12-05T13:28:26.197Z",
      "content": "<p>In order for the \"Submit to Competition\" button to be active after the Kernel commit, the following conditions must be met:</p>\n\n<p>CPU Kernel &lt;= 6 hours run-time\nGPU Kernel &lt;= 2 hours run-time\nNo internet access enabled\nNo multiple data sources enabled\nNo custom packages\nSubmission file must be named \"submission.csv\"</p>\n\n<p>Obviously there is a trick to go around some of this rules, simply copy pasting everything to the notebook and there wouldn't be a need to include data sources, and as for the run time requirements, we could always copy paste results in another kernel-&gt;commit and the \"Submit to Competition\" button will be active for you to try out your submission.csv .</p>\n\n<p>So how strict are the rules for submission before and when we reach stage two?  <a href=\"/inversion\">@inversion</a></p>",
      "rawMarkdown": "In order for the \"Submit to Competition\" button to be active after the Kernel commit, the following conditions must be met:\n\nCPU Kernel &lt;= 6 hours run-time\nGPU Kernel &lt;= 2 hours run-time\nNo internet access enabled\nNo multiple data sources enabled\nNo custom packages\nSubmission file must be named \"submission.csv\"\n\nObviously there is a trick to go around some of this rules, simply copy pasting everything to the notebook and there wouldn't be a need to include data sources, and as for the run time requirements, we could always copy paste results in another kernel-&gt;commit and the \"Submit to Competition\" button will be active for you to try out your submission.csv .\n\nSo how strict are the rules for submission before and when we reach stage two?  @inversion",
      "votes": 1,
      "replies": [
        {
          "id": 433896,
          "postDate": "2018-12-05T16:15:35.510Z",
          "content": "<p>If the intent of the rules are bypassed, the submitting team risks being removed from the leaderboard.</p>",
          "rawMarkdown": "If the intent of the rules are bypassed, the submitting team risks being removed from the leaderboard.",
          "votes": 1
        },
        {
          "id": 437212,
          "postDate": "2018-12-11T14:55:33.823Z",
          "content": "<p>I get:\n<code>\nYour runtime of ... minutes exceeds the CPU kernel max of 120 minutes\n</code>\nwhy?</p>",
          "rawMarkdown": "I get:\n```\nYour runtime of ... minutes exceeds the CPU kernel max of 120 minutes\n```\nwhy?"
        },
        {
          "id": 455995,
          "postDate": "2019-01-14T23:07:00.640Z",
          "content": "<p>becouse rules are : CPU Kernel &lt;= 6 hours run-time, GPU Kernel &lt;= 2 hours run-time</p>",
          "rawMarkdown": "becouse rules are : CPU Kernel &lt;= 6 hours run-time, GPU Kernel &lt;= 2 hours run-time",
          "votes": 1
        }
      ]
    },
    {
      "id": 426291,
      "postDate": "2018-11-23T02:33:08.757Z",
      "content": "<p>Can you give a quick summary of the stage-1/stage-2 logistics for those of us new to Kaggle?</p>",
      "rawMarkdown": "Can you give a quick summary of the stage-1/stage-2 logistics for those of us new to Kaggle?",
      "votes": 1,
      "replies": [
        {
          "id": 429460,
          "postDate": "2018-11-28T22:18:23.753Z",
          "content": "<p>Because this is a <strong>code competition</strong>, the \"two stage\" format simply means that on the final submission deadline (review the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification#Timeline\">Timeline</a> page for specifics), you must select your final submission(s) from the kernel(s) that will be run against the unseen private test set. \"Stage 1\" is effectively that final submission deadline, and \"Stage 2\" is the act of our platform running all submissions' code against the unseen private test set. This is specified in the last question on the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification#Kernels-FAQ\">Kernel FAQ</a>. The Stage 2 process will likely require 2 weeks of processing and verification time following final submission deadline.</p>",
          "rawMarkdown": "Because this is a **code competition**, the \"two stage\" format simply means that on the final submission deadline (review the [Timeline](https://www.kaggle.com/c/quora-insincere-questions-classification#Timeline) page for specifics), you must select your final submission(s) from the kernel(s) that will be run against the unseen private test set. \"Stage 1\" is effectively that final submission deadline, and \"Stage 2\" is the act of our platform running all submissions' code against the unseen private test set. This is specified in the last question on the [Kernel FAQ](https://www.kaggle.com/c/quora-insincere-questions-classification#Kernels-FAQ). The Stage 2 process will likely require 2 weeks of processing and verification time following final submission deadline.",
          "votes": 4
        },
        {
          "id": 439075,
          "postDate": "2018-12-14T17:32:50.123Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> We are experiencing different execution times for kernels and non-reproducable results with pytorch/keras. Will kaggle do multiple runs of kernels to average for stage 2 or just one time run. </p>\n\n<p>If its a onetime run I am wondering if CuDNN randomness could play a luck factor at top LB positions as our scores are very close. </p>",
          "rawMarkdown": "@juliaelliott We are experiencing different execution times for kernels and non-reproducable results with pytorch/keras. Will kaggle do multiple runs of kernels to average for stage 2 or just one time run. \n\nIf its a onetime run I am wondering if CuDNN randomness could play a luck factor at top LB positions as our scores are very close. "
        },
        {
          "id": 442766,
          "postDate": "2018-12-20T13:18:08.443Z",
          "content": "<p>Is there a time limit for the second phase?</p>",
          "rawMarkdown": "Is there a time limit for the second phase?"
        }
      ]
    },
    {
      "id": 425264,
      "postDate": "2018-11-21T11:09:52.363Z",
      "content": "<p>The rules state that no external data could be used. I just want to know that can we use pretrained weights which needs to download thorough the Internet for transfer learning? Thx.</p>",
      "rawMarkdown": "The rules state that no external data could be used. I just want to know that can we use pretrained weights which needs to download thorough the Internet for transfer learning? Thx.",
      "votes": 1,
      "replies": [
        {
          "id": 426126,
          "postDate": "2018-11-22T16:52:43.930Z",
          "content": "<p>There would be no way incorporate these within the competition constraints (i.e., if you create a kernel that uses an additional dataset or via kernel internet access, you won't be able to submit the kernel output). </p>",
          "rawMarkdown": "There would be no way incorporate these within the competition constraints (i.e., if you create a kernel that uses an additional dataset or via kernel internet access, you won't be able to submit the kernel output). ",
          "votes": 1
        }
      ]
    },
    {
      "id": 444139,
      "postDate": "2018-12-23T10:11:38.930Z",
      "content": "<p>Hi <a href=\"/inversion\">@inversion</a>,</p>\n\n<p>Does the GPU kernel in this competition using the latest Kaggle GPU image?</p>\n\n<p>I made a PR to add <code>cupy</code> and <code>pynvrtc</code>, and also got merged to master branch:\n<a href=\"https://github.com/Kaggle/docker-python/blob/master/gpu.Dockerfile#L57-L58\">https://github.com/Kaggle/docker-python/blob/master/gpu.Dockerfile#L57-L58</a></p>\n\n<p>But from the kernel it says that the package is not found, any suggestions?</p>\n\n<pre><code>ModuleNotFoundError: No module named 'cupy'\n</code></pre>",
      "rawMarkdown": "Hi @inversion,\n\nDoes the GPU kernel in this competition using the latest Kaggle GPU image?\n\nI made a PR to add `cupy` and `pynvrtc`, and also got merged to master branch:\nhttps://github.com/Kaggle/docker-python/blob/master/gpu.Dockerfile#L57-L58\n\nBut from the kernel it says that the package is not found, any suggestions?\n\n    ModuleNotFoundError: No module named 'cupy'\n    ",
      "votes": 2,
      "replies": [
        {
          "id": 455748,
          "postDate": "2019-01-14T13:38:21.300Z",
          "content": "<p><a href=\"/inversion\">@inversion</a> Could you give some reply to this problem? It is still not working. Thanks.</p>",
          "rawMarkdown": "@inversion Could you give some reply to this problem? It is still not working. Thanks."
        }
      ]
    },
    {
      "id": 443400,
      "postDate": "2018-12-21T14:57:50.663Z",
      "content": "<p>Hi, inversion, \nIn the data page it is mentioned : </p>\n\n<p>This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions.</p>\n\n<p>what does public leaderboard data remains the same for both versions mean?</p>\n\n<p>does this mean out of 376k examples in private part  56k examples will be same as the public leaderboard and 320k will be different?</p>",
      "rawMarkdown": "Hi, inversion, \nIn the data page it is mentioned : \n\nThis file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions.\n\nwhat does public leaderboard data remains the same for both versions mean?\n\ndoes this mean out of 376k examples in private part  56k examples will be same as the public leaderboard and 320k will be different?",
      "votes": 2
    },
    {
      "id": 424932,
      "postDate": "2018-11-20T22:39:42.040Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> can we get an answer on the use of pretrained weights? Would like to use transfer learning approach.</p>",
      "rawMarkdown": "@inversion can we get an answer on the use of pretrained weights? Would like to use transfer learning approach.",
      "votes": 2,
      "replies": [
        {
          "id": 426124,
          "postDate": "2018-11-22T16:51:11.847Z",
          "content": "<p>There would be no way incorporate these within the competition constraints (i.e., if you create a kernel that uses an additional dataset, you won't be able to submit the kernel output). The exception would be tools that are contained in the Kaggle Docker packages.</p>",
          "rawMarkdown": "There would be no way incorporate these within the competition constraints (i.e., if you create a kernel that uses an additional dataset, you won't be able to submit the kernel output). The exception would be tools that are contained in the Kaggle Docker packages.",
          "votes": 1
        }
      ]
    },
    {
      "id": 473598,
      "postDate": "2019-02-18T08:40:29.283Z",
      "content": "<p>Welcoe too.</p>",
      "rawMarkdown": "Welcoe too."
    },
    {
      "id": 464299,
      "postDate": "2019-01-31T15:01:35.580Z",
      "content": "<p>My first challenge here on Kaggle, hope to help you and learn new skills. :)</p>",
      "rawMarkdown": "My first challenge here on Kaggle, hope to help you and learn new skills. :)"
    },
    {
      "id": 462643,
      "postDate": "2019-01-28T16:29:16.690Z",
      "content": "<p><a href=\"/inversion\">@inversion</a>\nHI\nIt is allowed loading \"kernel output files\" from my previous Quora kernels (my work) as input files?\nbecause it seems that this prevents me to submit output   predictions file! ( \"Submit to competition\" button remains inactive)</p>\n\n<p>thanks</p>",
      "rawMarkdown": "@inversion\nHI\nIt is allowed loading \"kernel output files\" from my previous Quora kernels (my work) as input files?\nbecause it seems that this prevents me to submit output   predictions file! ( \"Submit to competition\" button remains inactive)\n\nthanks"
    },
    {
      "id": 462418,
      "postDate": "2019-01-28T09:13:00.013Z",
      "content": "<p>Can single participants safely ignore the merger deadline?  No team, no worries?</p>",
      "rawMarkdown": "Can single participants safely ignore the merger deadline?  No team, no worries?"
    },
    {
      "id": 461642,
      "postDate": "2019-01-26T16:16:59.567Z",
      "content": "<p><a href=\"/inversion\">@inversion</a>\nHow can I set num_workers for multiprocessing for stage2? \nwhen I tested it in kernel, cpu_core isn't equal.</p>",
      "rawMarkdown": "@inversion\nHow can I set num_workers for multiprocessing for stage2? \nwhen I tested it in kernel, cpu_core isn't equal.",
      "replies": [
        {
          "id": 461662,
          "postDate": "2019-01-26T16:53:37.183Z",
          "content": "<p>What values are you getting?</p>",
          "rawMarkdown": "What values are you getting?"
        },
        {
          "id": 461758,
          "postDate": "2019-01-26T23:47:07.507Z",
          "content": "<p>number of cpu cores, \nbut it isn't fixed (2 or 4)</p>",
          "rawMarkdown": "number of cpu cores, \nbut it isn't fixed (2 or 4)"
        }
      ]
    },
    {
      "id": 461590,
      "postDate": "2019-01-26T13:03:26.553Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> I am trying to build a model using the Google Pre Trained Word Embeddings but my kernel dies just at the end! I don't understand what's happening!</p>",
      "rawMarkdown": "@inversion I am trying to build a model using the Google Pre Trained Word Embeddings but my kernel dies just at the end! I don't understand what's happening!",
      "replies": [
        {
          "id": 461660,
          "postDate": "2019-01-26T16:53:12.157Z",
          "content": "<p>There have been some operational issues with kernels, but should be resolved now. Can you try again and let me know what happens? Thx.</p>",
          "rawMarkdown": "There have been some operational issues with kernels, but should be resolved now. Can you try again and let me know what happens? Thx."
        }
      ]
    },
    {
      "id": 457529,
      "postDate": "2019-01-17T16:20:23.017Z",
      "content": "<p>HI\nIs WordNet synsets allowed in this competition?\nAs I can run wordnet synset using Kaggle kernel, I will assume its OK, just want to double check.\nThanks,</p>",
      "rawMarkdown": "HI\nIs WordNet synsets allowed in this competition?\nAs I can run wordnet synset using Kaggle kernel, I will assume its OK, just want to double check.\nThanks,"
    },
    {
      "id": 456831,
      "postDate": "2019-01-16T16:04:51.360Z",
      "content": "<p>When clicking 'submit to competition', it always redirect to 404 page. Could anyone help?</p>\n\n<p>Update: The reason is I saved submission file to different location like <code>sub.to_csv(\"experiment_folder/submission.csv\", index=False)</code>. I think we must do <code>sub.to_csv(\"submission.csv\", index=False)</code>, and I finally submitted successfully.</p>",
      "rawMarkdown": "When clicking 'submit to competition', it always redirect to 404 page. Could anyone help?\n\nUpdate: The reason is I saved submission file to different location like `sub.to_csv(\"experiment_folder/submission.csv\", index=False)`. I think we must do `sub.to_csv(\"submission.csv\", index=False)`, and I finally submitted successfully.",
      "replies": [
        {
          "id": 456832,
          "postDate": "2019-01-16T16:08:01.153Z",
          "content": "<p><a href=\"/inversion\">@inversion</a></p>",
          "rawMarkdown": "@inversion"
        },
        {
          "id": 456849,
          "postDate": "2019-01-16T16:46:51.847Z",
          "content": "<p>What happens if you create a brand new kernel and just read/submit the sample submission?</p>",
          "rawMarkdown": "What happens if you create a brand new kernel and just read/submit the sample submission?"
        },
        {
          "id": 457125,
          "postDate": "2019-01-17T01:31:33.397Z",
          "content": "<p>Thanks for reply. I succeeded to submit sample submission.</p>",
          "rawMarkdown": "Thanks for reply. I succeeded to submit sample submission.",
          "votes": 1
        }
      ]
    },
    {
      "id": 453076,
      "postDate": "2019-01-09T16:26:03.930Z",
      "content": "<p>Is data augmentation ( by synonyms , from embedding ) is allowed in this competition or will it be considered as external data ?</p>",
      "rawMarkdown": "Is data augmentation ( by synonyms , from embedding ) is allowed in this competition or will it be considered as external data ?"
    },
    {
      "id": 451935,
      "postDate": "2019-01-07T23:24:18.410Z",
      "content": "<p>Could you please install GPU version of Apache MXNet? Currently, only CPU version is installed, and if I try to use GPU I get an error. Here is  the example:\n<code>\nimport mxnet as mx\na = mx.nd.array([1,2,34], mx.gpu(0))\n</code>\nThe error message is:</p>\n\n<p><code>\nMXNetError: [23:21:15] src/storage/storage.cc:137: Compile with USE_CUDA=1 to enable GPU usage Stack trace returned 10 entries: [bt] (0) /opt/conda/lib/python3.6/site-packages/mxnet/libmxnet.so(+0x1e945a) [0x7f45bb45445a] [bt] (1) /opt/conda/lib/python3.6/site-packages/mxnet/libmxnet.so(+0x1e9ac1) [0x7f45bb454ac1] [bt] (2) /opt/conda/lib/python3.6/site-packages/mxnet/libmxnet.so(+0x2f83205) [0x7f45be1ee205]\n</code></p>\n\n<p>To do so, MXNet should be installed with <code>pip install mxnet-cu90mkl</code> (for CUDA 9.0 and MKL support).\nThank you.</p>",
      "rawMarkdown": "Could you please install GPU version of Apache MXNet? Currently, only CPU version is installed, and if I try to use GPU I get an error. Here is  the example:\n```\nimport mxnet as mx\na = mx.nd.array([1,2,34], mx.gpu(0))\n```\nThe error message is:\n\n```\nMXNetError: [23:21:15] src/storage/storage.cc:137: Compile with USE_CUDA=1 to enable GPU usage Stack trace returned 10 entries: [bt] (0) /opt/conda/lib/python3.6/site-packages/mxnet/libmxnet.so(+0x1e945a) [0x7f45bb45445a] [bt] (1) /opt/conda/lib/python3.6/site-packages/mxnet/libmxnet.so(+0x1e9ac1) [0x7f45bb454ac1] [bt] (2) /opt/conda/lib/python3.6/site-packages/mxnet/libmxnet.so(+0x2f83205) [0x7f45be1ee205]\n```\n\nTo do so, MXNet should be installed with `pip install mxnet-cu90mkl` (for CUDA 9.0 and MKL support).\nThank you."
    },
    {
      "id": 451932,
      "postDate": "2019-01-07T23:17:43.210Z",
      "content": "<p>Could you please allow usage of <a href=\"http://gluon-nlp.mxnet.io/\">http://gluon-nlp.mxnet.io/</a> package? It is state of the art NLP library for Apache MXNet framework, and it is as easy to install as <code>pip install gluonnlp</code>. Right now I can import it for CPU instance only.</p>",
      "rawMarkdown": "Could you please allow usage of http://gluon-nlp.mxnet.io/ package? It is state of the art NLP library for Apache MXNet framework, and it is as easy to install as `pip install gluonnlp`. Right now I can import it for CPU instance only."
    },
    {
      "id": 451357,
      "postDate": "2019-01-07T00:03:34.540Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> the training set looks like it contains quite a few false negatives (which is pretty understandable!). Can you tell us if the test set labelled in the same way or has it been human validated?\nThanks for setting this up.</p>",
      "rawMarkdown": "@inversion the training set looks like it contains quite a few false negatives (which is pretty understandable!). Can you tell us if the test set labelled in the same way or has it been human validated?\nThanks for setting this up."
    },
    {
      "id": 450254,
      "postDate": "2019-01-04T15:15:22.533Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> is it fair to include pseudo labeling manually ? i.e. print the prediction then add it again to the training using copy paste ?</p>",
      "rawMarkdown": "@inversion is it fair to include pseudo labeling manually ? i.e. print the prediction then add it again to the training using copy paste ?"
    },
    {
      "id": 449977,
      "postDate": "2019-01-04T03:59:45.917Z",
      "content": "<p>HI can someone help me with an error. Its weird that I can't post my problems in discussions as well. When I try to submit my output, it says that I cannot use docker image. What does it mean?</p>",
      "rawMarkdown": "HI can someone help me with an error. Its weird that I can't post my problems in discussions as well. When I try to submit my output, it says that I cannot use docker image. What does it mean?\n\n"
    },
    {
      "id": 446681,
      "postDate": "2018-12-28T13:39:35.507Z",
      "content": "<p>i am a new nlper! quora  competition is very funny!</p>",
      "rawMarkdown": "i am a new nlper! quora  competition is very funny!"
    },
    {
      "id": 446019,
      "postDate": "2018-12-27T11:01:58.260Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> \nI find an interesting thing, the distribution between training data and testing data is different.\nSo I want to know whether the public and private testing data are derived from a large test set, then randomly cut into two parts, or derived from two data sets? </p>",
      "rawMarkdown": "@inversion \nI find an interesting thing, the distribution between training data and testing data is different.\nSo I want to know whether the public and private testing data are derived from a large test set, then randomly cut into two parts, or derived from two data sets? "
    },
    {
      "id": 444396,
      "postDate": "2018-12-24T00:11:58.493Z",
      "content": "<p>Hey, This is Chandrashekhar Adhage...Excited for this Competition..:)</p>",
      "rawMarkdown": "Hey, This is Chandrashekhar Adhage...Excited for this Competition..:)"
    },
    {
      "id": 442557,
      "postDate": "2018-12-20T05:51:16.840Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> Are the questions in languages other than english correctly labeled (insecure or not)? When labeling the training data, were different models used for different languages? </p>",
      "rawMarkdown": "@inversion Are the questions in languages other than english correctly labeled (insecure or not)? When labeling the training data, were different models used for different languages? "
    },
    {
      "id": 441155,
      "postDate": "2018-12-18T10:34:40.537Z",
      "content": "<p>Is it fair to manually remove or remark false positives of the dataset inside the code of the kernel?\nfor example:\n<code>\ndata[data['question_text'] in [TEXTS TO MANUALLY MARK]]['target']=0\n</code></p>",
      "rawMarkdown": "Is it fair to manually remove or remark false positives of the dataset inside the code of the kernel?\nfor example:\n```\ndata[data['question_text'] in [TEXTS TO MANUALLY MARK]]['target']=0\n```",
      "replies": [
        {
          "id": 443858,
          "postDate": "2018-12-22T15:50:43.117Z",
          "content": "<p>Yes.</p>",
          "rawMarkdown": "Yes."
        }
      ]
    },
    {
      "id": 440682,
      "postDate": "2018-12-17T22:08:42.677Z",
      "content": "<p>Hi, this is my first Kaggle competition. I submitted and got a score of 0.37. Does this mean I got a 37% classification accuracy?</p>",
      "rawMarkdown": "Hi, this is my first Kaggle competition. I submitted and got a score of 0.37. Does this mean I got a 37% classification accuracy?",
      "replies": [
        {
          "id": 440830,
          "postDate": "2018-12-18T02:21:16.650Z",
          "content": "<p>The score in public LB has calculated with F1, not accuracy. you can see it in <a href=\"https://en.wikipedia.org/wiki/F1_score\">https://en.wikipedia.org/wiki/F1_score</a></p>",
          "rawMarkdown": "The score in public LB has calculated with F1, not accuracy. you can see it in https://en.wikipedia.org/wiki/F1_score",
          "votes": 1
        },
        {
          "id": 441631,
          "postDate": "2018-12-18T21:38:54.340Z",
          "content": "<p>I see, thanks.</p>",
          "rawMarkdown": "I see, thanks."
        }
      ]
    },
    {
      "id": 440329,
      "postDate": "2018-12-17T12:09:00.173Z",
      "content": "<p>In the Private LB phase, will our model be re-run? Or use the parameters of the current optimal model to directly predict the results of the new test data?</p>",
      "rawMarkdown": "In the Private LB phase, will our model be re-run? Or use the parameters of the current optimal model to directly predict the results of the new test data?",
      "replies": [
        {
          "id": 442503,
          "postDate": "2018-12-20T03:11:27.540Z",
          "content": "<p>It will be re-run. <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/70715#440595\">here</a></p>",
          "rawMarkdown": "It will be re-run. [here](https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/70715#440595)"
        },
        {
          "id": 442777,
          "postDate": "2018-12-20T13:40:39.937Z",
          "content": "<p>3q</p>",
          "rawMarkdown": "3q"
        }
      ]
    },
    {
      "id": 435382,
      "postDate": "2018-12-07T23:44:35.633Z",
      "content": "<p>Considering the fact, that the Kernel FAQ states \"No multiple data sources enabled\": Does this apply to the given word vector files? Meaning, are we only allowed to use one of those word vector files? </p>\n\n<p>And if I create a dict object in the kernel which contains words to be replaced, does this count as a data source? This question pertains to any form of data manipulation of the training/test data which is based on a manual analysis of this data and results in some form of data collection.</p>\n\n<p>And if such analysis data - like the dict above - counts as a data source, does this lead to a conflict with loaded word vector files - meaning only the dict or the word vectors are allowed?</p>",
      "rawMarkdown": "Considering the fact, that the Kernel FAQ states \"No multiple data sources enabled\": Does this apply to the given word vector files? Meaning, are we only allowed to use one of those word vector files? \n\nAnd if I create a dict object in the kernel which contains words to be replaced, does this count as a data source? This question pertains to any form of data manipulation of the training/test data which is based on a manual analysis of this data and results in some form of data collection.\n\nAnd if such analysis data - like the dict above - counts as a data source, does this lead to a conflict with loaded word vector files - meaning only the dict or the word vectors are allowed?",
      "replies": [
        {
          "id": 444861,
          "postDate": "2018-12-25T02:15:54.603Z",
          "content": "<p>Don't sweat it. Unless your dict is like 5k,10k+ and clearly not hand built, they won't dock you for that. That's considered part of your pre-processing.</p>",
          "rawMarkdown": "Don't sweat it. Unless your dict is like 5k,10k+ and clearly not hand built, they won't dock you for that. That's considered part of your pre-processing."
        }
      ]
    },
    {
      "id": 434910,
      "postDate": "2018-12-07T06:22:45.743Z",
      "content": "<p>can i train different models on different kernels, and combine these pretrained models generated on final kernel?</p>",
      "rawMarkdown": "can i train different models on different kernels, and combine these pretrained models generated on final kernel?"
    },
    {
      "id": 431921,
      "postDate": "2018-12-03T05:15:52.257Z",
      "content": "<p>I recently started working on ML problems as became more aware of ML.Net (Microsoft's open source ML library). Is there anyway i can get score for my submission which is in Output folder but was uploaded from local drive since I had worked on this locally. I understand that I did not adhere to rules, and I am fine with being disqualified. But I would love to see what was the score for my predictions.</p>",
      "rawMarkdown": "I recently started working on ML problems as became more aware of ML.Net (Microsoft's open source ML library). Is there anyway i can get score for my submission which is in Output folder but was uploaded from local drive since I had worked on this locally. I understand that I did not adhere to rules, and I am fine with being disqualified. But I would love to see what was the score for my predictions.",
      "replies": [
        {
          "id": 444873,
          "postDate": "2018-12-25T03:00:38.500Z",
          "content": "<p>Write a script to serialize your output as a (long) sequence of assignment statements to (e.g.) a numpy array.  You can then cut&amp;paste that into a script and have it regurgitate your offline results.  Won't be a valid competition entry, but should enable you to get a metric for your personal usage.</p>",
          "rawMarkdown": "Write a script to serialize your output as a (long) sequence of assignment statements to (e.g.) a numpy array.  You can then cut&amp;paste that into a script and have it regurgitate your offline results.  Won't be a valid competition entry, but should enable you to get a metric for your personal usage."
        }
      ]
    },
    {
      "id": 429367,
      "postDate": "2018-11-28T18:43:50.367Z",
      "content": "<p>Hi. GloVe: Global Vectors for Word Representation using nltk.tokenize.stanford.StanfordTokenizer, but we can't use it without stanford-postagger.jar (3,49 Mb) and os.environ['JAVAHOME'] = java_path. Can you add this file to data?\nZip with file: <a href=\"https://nlp.stanford.edu/software/stanford-postagger-2018-10-16.zip\">https://nlp.stanford.edu/software/stanford-postagger-2018-10-16.zip</a></p>",
      "rawMarkdown": "Hi. GloVe: Global Vectors for Word Representation using nltk.tokenize.stanford.StanfordTokenizer, but we can't use it without stanford-postagger.jar (3,49 Mb) and os.environ['JAVAHOME'] = java_path. Can you add this file to data?\nZip with file: https://nlp.stanford.edu/software/stanford-postagger-2018-10-16.zip"
    },
    {
      "id": 427359,
      "postDate": "2018-11-25T10:20:30.683Z",
      "content": "<p>Could you please update the python package like \"fastai\",  thanks :)</p>",
      "rawMarkdown": "Could you please update the python package like \"fastai\",  thanks :)"
    },
    {
      "id": 426826,
      "postDate": "2018-11-24T00:17:44.563Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> I am curious to know how exactly the overall F1 score is being calculated here? Is it micro, macro, or some other?</p>",
      "rawMarkdown": "@inversion I am curious to know how exactly the overall F1 score is being calculated here? Is it micro, macro, or some other?",
      "replies": [
        {
          "id": 450636,
          "postDate": "2019-01-05T12:42:45.573Z",
          "content": "<p>No differences since it's a binary classification, only two classes.</p>",
          "rawMarkdown": "No differences since it's a binary classification, only two classes."
        }
      ]
    },
    {
      "id": 426232,
      "postDate": "2018-11-22T23:46:47.633Z",
      "content": "<p>Insincere questions are unsafe.</p>",
      "rawMarkdown": "Insincere questions are unsafe."
    },
    {
      "id": 422090,
      "postDate": "2018-11-15T19:07:57.437Z",
      "content": "<p>Can I use the information on the test dataset as part of my training? Like to extract all the words from it to the model's vocabulary. Thanks.</p>",
      "rawMarkdown": "Can I use the information on the test dataset as part of my training? Like to extract all the words from it to the model's vocabulary. Thanks."
    },
    {
      "id": 419442,
      "postDate": "2018-11-12T01:04:35.627Z",
      "content": "<p>Intresting NLP-related competition on Kaggle! </p>",
      "rawMarkdown": "Intresting NLP-related competition on Kaggle! \n\n"
    },
    {
      "id": 417514,
      "postDate": "2018-11-08T12:01:26.420Z",
      "content": "<p>Am i eligible to use embeddings which trained on competition data?</p>",
      "rawMarkdown": "Am i eligible to use embeddings which trained on competition data?",
      "replies": [
        {
          "id": 417913,
          "postDate": "2018-11-09T01:07:56.013Z",
          "content": "<p>Yes. You can utilize any data that is generated during the Kernel run.</p>",
          "rawMarkdown": "Yes. You can utilize any data that is generated during the Kernel run."
        }
      ]
    },
    {
      "id": 450320,
      "postDate": "2019-01-04T17:18:27.700Z",
      "content": "<p><a href=\"/inversion\">@inversion</a> I have a question that worries me a bit:\nAbove private LB it says \"The private leaderboard is calculated over the same rows as the public leaderboard in this competition.\"</p>\n\n<p>It is now clear that kernels will be re-run in phase 2 to predict labels on all 376k rows.\nBut then to get f1 score there are two options:\n(1) 376k rows are partitioned into 56k and 320k to get f1 score of public and private LB respectively.\n(2) 376k row are reorganized into 56k and 376k to get f1 score of public and private LB respectively. </p>\n\n<p>Which of (1) or (2) will be the procedure?</p>",
      "rawMarkdown": "@inversion I have a question that worries me a bit:\nAbove private LB it says \"The private leaderboard is calculated over the same rows as the public leaderboard in this competition.\"\n\nIt is now clear that kernels will be re-run in phase 2 to predict labels on all 376k rows.\nBut then to get f1 score there are two options:\n(1) 376k rows are partitioned into 56k and 320k to get f1 score of public and private LB respectively.\n(2) 376k row are reorganized into 56k and 376k to get f1 score of public and private LB respectively. \n\nWhich of (1) or (2) will be the procedure?",
      "isDeleted": true
    },
    {
      "id": 450305,
      "postDate": "2019-01-04T16:52:53.470Z",
      "content": "<p>In description it says test data has ~56k rows in stage 1 and ~376k rows in stage 2. \nThus, ~320k extra rows will be included for stage 2.</p>\n\n<p>Are these 56k and 320k rows random samples of the full test dataset of 376k rows?\n... or are they two separate datasets, possibly with a very different frequency of class 1 (insincere questions) ?</p>",
      "rawMarkdown": "In description it says test data has ~56k rows in stage 1 and ~376k rows in stage 2. \nThus, ~320k extra rows will be included for stage 2.\n\nAre these 56k and 320k rows random samples of the full test dataset of 376k rows?\n... or are they two separate datasets, possibly with a very different frequency of class 1 (insincere questions) ?",
      "isDeleted": true
    },
    {
      "id": 418573,
      "postDate": "2018-11-10T06:52:47.123Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 418023,
      "postDate": "2018-11-09T05:51:20.217Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 417168,
      "postDate": "2018-11-07T21:58:01.290Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 473458,
      "postDate": "2019-02-18T03:08:32.027Z",
      "content": "<p>Thanks a lot</p>",
      "rawMarkdown": "Thanks a lot"
    }
  ],
  "comments": [
    {
      "id": 417390,
      "author_name": "IridiumBlue",
      "author_url": "",
      "post_date": "2018-11-08T07:20:13.583000",
      "content": "<p>Looking at the training data, am I the only one who sees lots of entries which seem to have been tagged w/o careful review?  Examples - </p>\n\n<p>0098e5b8335a62791abf,Is Dr. Steven Greer an extremely good fraudster with his claims of ufo' s and aliens?,1\n009afb0a3716d819bf7e,What did Trump and Sessions possibly gain from firing McCabe with less than 24 hours until retiring?,1\nWhy do progressive activists feel that tearing down antique historic but politically incorrect statues accomplishes a positive good?,1\ndc002c8cff10ddc797dc,Does Trump's inaction on his own justice department's indictment of Russian cyber attacks mean he has violated his oath of office?,1\ndc16793a4d38a9a27803,How do black people celebrate Martin Luther King Jr's birthday?,1\n00b92227d01a117d9c85,\"Do white people in America think they are privileged? Is that a real thing going on, or just a bad meme?\",1\ndbd98d3003044f51791a,Do Democrats have a plan to curb violence in Chicago?,1\n00c6da623540ba7d0a9d,\"What would have to be proven to show that Cohen is guilty of influence peddling, a crime? Any more than is already a matter of public record?\",1\ndc44f81813106a237458,Why doesn't China apologize for invading Tibet? They did it even before the Sino-Indian War.,1\n00c7fbcb0495b64ed7c2,\"If Donald Trump pardons himself, would that count as an admission of guilt and make it easier for the state attorney general to prosecute on related but not identical charges?\",1\n00cc8173e3c8a2a5c3fd,Can asperger's recognise sarcasm in relationships?,1 \n00df0e17c0de8d326227Do some atheists feel guilty about indoctrinating their children with atheism and not giving them hope and a moral compass?,1\ndcb2c4b6be697b16c9e3,How do I avoid losing the argument when I try to defend ethnic groups with statistically proven higher criminality rates? Am I wrong to defend these ethnic groups in the face of hard facts?,1\ndcbaf4d21e0b29592e8c,Does the conservatives’ view on climate change contradict their core conservative values?,1\ndc972c2e4f084dc8fdb8,\"Why do so many Americans get puzzled when I ask for a serviette in their restaurants, but know right away what I am talking about when I ask for a Napkin (aka diaper)?\",1\n00df34daff54eef050f7What is it like to be a gay in IIT Madras?,1</p>\n\n<p>But here we have  : dbae47ac9a7adf7f178cWhat is it like to be gay in Southern California?,0</p>\n\n<p>(So - gay as an adjective is OK, but not as a noun?)</p>\n\n<p>This is about 10 % of the 1's, near as I can reckon.    Yes, these are ham-fisted questions about delicate topics.   Yes, they could be worded more gently.   But they do not harass nor disparage, they certainly appear to be in good faith and seeking real answers.</p>\n\n<p>There's a real garbage-in garbage-out problem here.   The dataset appears not to have been human-curated at all, but rather from an algorithm which is surfacing false positives.   Or from a curator who is all nerve endings about certain issues.</p>\n\n<p>I was very excited to enter (and win!!!) this competition but I am stopped dead in my tracks by bad training data.  My complaint isn't about politics, I could not describe to a human with a PhD in English Literature how to arrive at the conclusions given in this sample.</p>\n\n<p>Trying to code it is a fool's errand, to put it mildly.</p>\n\n<p>Not convinced?   How about this one - </p>\n\n<p>dcbfe10facdd66c91eca,Why aren't there any women in the Navy SEALS?,1</p>\n\n<p>Explain that one.   To the US Navy.    <a href=\"https://www.navytimes.com/news/your-navy/2018/02/16/two-women-could-enter-navy-special-operations-training-this-year/\">https://www.navytimes.com/news/your-navy/2018/02/16/two-women-could-enter-navy-special-operations-training-this-year/</a></p>",
      "votes": 22,
      "replies": [
        {
          "id": 417641,
          "author_name": "Quentin Retourne",
          "author_url": "",
          "post_date": "2018-11-08T15:07:04.427000",
          "content": "<p>I feel like some of the questions you highlighted are biased or too badly-worded, thus flagged as insincere.</p>\n\n<p>Yet I have to recognize that some questions should not be highlighted as insincere (a gay vs. gay is the best example).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 417804,
          "author_name": "Anne Clark",
          "author_url": "",
          "post_date": "2018-11-08T19:53:43.393000",
          "content": "<p>I had the same concerns after looking through some of the questions that were flagged/not flagged. You focus on things that seem incorrectly flagged, but there are also a lot of things they've obviously missed flagging like single word/nonsense kinds of things. I'm concerned we're not going to be able to train anything better than what they've already got when the training data is so poorly curated. I was also wondering whether the data set we're ultimately tested on will be of similar quality.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 417819,
          "author_name": "IridiumBlue",
          "author_url": "",
          "post_date": "2018-11-08T20:24:42.973000",
          "content": "<p>I hadn't even looked at the false negatives.   I suppose if the final data testing data is of low quality, that should lower everyone's score by a similar amount, so the final ranking may still be accurate.  But maybe not.</p>\n\n<p>Are the false negatives all simply malformed/meaningless?  I don't see that the target is defined to include those (<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/data\">https://www.kaggle.com/c/quora-insincere-questions-classification/data</a>)</p>\n\n<p>But too much noise in the training data - yea, can't really progress much with garbage-in.</p>\n\n<p>Should we all gang up and try to clean up the data?   There are 100K flagged records, it takes about 2 seconds to check each, if 50 people spend an hour each, and then another 1/2 hour to check each other's work - that would do it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 417902,
          "author_name": "Anne Clark",
          "author_url": "",
          "post_date": "2018-11-09T00:38:37.040000",
          "content": "<p>Right, it's still fair if we've all got the same data, but probably less fun and useful.</p>\n\n<p>And you're right that the description didn't specifically include malformed/meaningless things in the insincere category. But if you look at the really short questions, for example, it looks like they are actually trying to flag malformed questions:</p>\n\n<p>sincere: 'Is God 42?', 'In Islam?', 'I 12?', 'How can I?', 'Am I fake?', 'Wat is 1A?', 'What sexy?', 'Why is 16?', 'What meow?', 'Why is αθ?', 'Hello sir?', 'IS 1+1 21?', 'ESR is 35?'</p>\n\n<p>insincere: 'Quora: ', 'Is ', 'Dowry: ', 'If ', 'To Quora:', 'Sexism: ', 'Hungary: ', 'History: ', 'Islam: ', 'Politics: ', 'Feminism: ', '\"', 'Germany: ', 'Nepal: ', 'India: ', \"I'm an \", 'Wicca: ', 'Bye Bye?', 'Incest: ', 'Do ', 'Poland: ', 'W', 'Why is ', 'Identity: '</p>\n\n<p>Anyway, I don't really have a sense at this point of whether the false positives or false negatives are a bigger issue overall.  </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 420230,
          "author_name": "Vishy",
          "author_url": "",
          "post_date": "2018-11-13T10:14:16.467000",
          "content": "<p>I have got a few examples which I believe should be flagged, but i guess those flagged are sufficient to be used for training now.</p>\n\n<p>0015a0bfd83e1fa30b73 - Which black superhero can beat Superman or is as strong as him?\n00c25e8c423fa71efcb2 - Why does this girl of a different race than me think I am racist just becausedo not find her physically attractive?\n0089e5eb91ca56c62a8b - Should I be concerned if I saw a white person with a black child?</p>\n\n<p>Pleas share your thoughts from <a href=\"/inversion\">@inversion</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 430132,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-11-29T22:02:46.150000",
          "content": "<p>@IridiumBlue: So - gay as an adjective is OK, but not as a noun?</p>\n\n<p>That seems perfectly reasonable to me. In my experience, in English, \"gay\" is normally an adjective (though it may occasionally be used as a noun), but in Trollspeak, it is commonly used as a noun (and maybe not often as an adjective, except in phrases like \"that's so gay\").</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 430324,
          "author_name": "Calisthenics",
          "author_url": "",
          "post_date": "2018-11-30T07:56:11.260000",
          "content": "<p>This is indeed a serious problem, quite a lot of the questions are false positive/ false negatives, which makes it quite difficult to train an efficient algorithm. We need to have a better dataset, otherwise our algorithms will not be so useful, and only be approximating the way the Quora team( combining  algorithms and manual labeling) label the insincere questions, not so practical then.\nI wish they can make an effort trying to clean some data.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 422736,
      "author_name": "Shujian Liu",
      "author_url": "",
      "post_date": "2018-11-16T17:42:23.147000",
      "content": "<p>Seems most of us use CUDNN package which has strong randomness for each run (but very fast). Is that possible in stage 2, you run the model a few times and take an average?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 439077,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-14T17:34:39.127000",
          "content": "<p>Please answer this <a href=\"/inversion\">@inversion</a> , <a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>We are experiencing huge randomness and LB is not very reliable </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 439106,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-14T18:40:27.987000",
          "content": "<p>It would be unfair for those though who try to tune their model to be as stable as possible.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 440595,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-12-17T18:40:21.487000",
          "content": "<p>Kernels will only be run once.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 441637,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-18T21:50:38.533000",
          "content": "<p>Thanks <a href=\"/inversion\">@inversion</a>, makes total sense to me. One more question though: As the test size for private LB will be larger, the kernel will need more time to run. Do we need to make an estimate for that, or is it enough if the current public LB version runs within 2 hours?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 441672,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-12-18T23:09:30.007000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Kernels will need to take into consideration the additional processing and inference time of the larger Test dataset. That will need to fit into the 2 / 6 hour constraints.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 416988,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2018-11-07T15:13:24.570000",
      "content": "<p>Great to have another NLP-related competition on Kaggle! Surely going to give this a go!</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 416747,
      "author_name": "Toby Cheese",
      "author_url": "",
      "post_date": "2018-11-07T08:04:31.293000",
      "content": "<p>Can the docker image change during the course of the competition? E.g. can it happen that some interesting library will not be available at the beginning, but and made available later? If so, do I have to monitor commits to the official docker image to make myself aware of this?</p>",
      "votes": 8,
      "replies": [
        {
          "id": 417904,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-09T00:53:22.167000",
          "content": "<p>Great question! We expect the Kaggle Docker image to have updates over the course of the competition, and there are no restrictions to using the latest updates.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 417956,
          "author_name": "bilal2vec",
          "author_url": "",
          "post_date": "2018-11-09T03:49:47.550000",
          "content": "<p>How long does it usually take for a pull request to the docker-python repository to be added to the kernel's image?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 416527,
      "author_name": "MichaelP",
      "author_url": "",
      "post_date": "2018-11-06T20:14:01.027000",
      "content": "<p>In stage 2 of the competition you rerun the kernels on a test set of approx 376k vs approx 56k in the first stage. Is there any adjustment made to the available run times in the kernels, or do they still have to run in 6 hours (or 2 hours for GPU)?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 416694,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2018-11-07T05:34:50.440000",
          "content": "<p>The run times (as you’ve stated) will remain the same.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 417190,
          "author_name": "MichaelP",
          "author_url": "",
          "post_date": "2018-11-07T23:11:47.253000",
          "content": "<p>Thanks for confirming.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437209,
          "author_name": "berner",
          "author_url": "",
          "post_date": "2018-12-11T14:49:45.260000",
          "content": "<p><code>\nyour runtime of ... minutes exceeds the CPU kernel max of 120 minutes\n</code>\nwhat exactly happened here?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 424161,
      "author_name": "Mohsin hasan",
      "author_url": "",
      "post_date": "2018-11-19T17:04:03.737000",
      "content": "<p><a href=\"/inversion\">@inversion</a> Are vectors and other models from spaCy fair game??</p>\n\n<p>Other people have asked same thing here: <a href=\"https://www.kaggle.com/jpmiller/bonus-vectors-with-spacy/comments\">https://www.kaggle.com/jpmiller/bonus-vectors-with-spacy/comments</a></p>\n\n<p>Thanks!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 424347,
          "author_name": "Sava Kalbachou",
          "author_url": "",
          "post_date": "2018-11-20T00:54:50.357000",
          "content": "<p><a href=\"/tezdhar\">@tezdhar</a> I'd extend this question even more: Are vectors and other models from available Docker image fair game?</p>\n\n<p><a href=\"/inversion\">@inversion</a> sorry to ping you again, but please answer on these questions asap. It just might save a lot of time for all participants.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 426039,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-11-22T13:45:28.560000",
          "content": "<p><a href=\"/inversion\">@inversion</a> Sorry to also ping you again. But a reply to that question is pretty crucial.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 426119,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-22T16:48:50.820000",
          "content": "<p>Tools that are available in the Docker image are fair game.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 443744,
          "author_name": "Abed Khooli",
          "author_url": "",
          "post_date": "2018-12-22T10:11:28.020000",
          "content": "<p>Now that the current docker has fastai v1 which includes the WT103 (wiki text) language model (not among white listed), is this answer still valid?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 443768,
          "author_name": "Sava Kalbachou",
          "author_url": "",
          "post_date": "2018-12-22T11:36:06.607000",
          "content": "<p>Hi <a href=\"/abedkhooli\">@abedkhooli</a>,\nAre you sure that it's possible to use fasta1 pre-trained LM without internet access? Shouldn't it be downloaded separately from the package?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 443800,
          "author_name": "Abed Khooli",
          "author_url": "",
          "post_date": "2018-12-22T13:10:57.763000",
          "content": "<p>Technically speaking, the model gets downloaded (at run time) if referenced in the kernel code but need not add as part of the datasets, that's why I asked for a definite answer from <a href=\"/inversion\">@inversion</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 443805,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2018-12-22T13:33:50.283000",
          "content": "<p>You cannot access Internet through runtime in scoring afaik.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 445599,
          "author_name": "Abed Khooli",
          "author_url": "",
          "post_date": "2018-12-26T19:39:54.050000",
          "content": "<p>It runs but after commit the submission link is disabled. The tooltip reads:\n\"You cannot use internet access for this competition. Your runtime of 134 minutes exceeds the GPU kernel max of 120 minutes.\"\nSo, apparently there is limit on tools and compute.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 417221,
      "author_name": "bangda",
      "author_url": "",
      "post_date": "2018-11-08T00:44:49.207000",
      "content": "<p>Hi inversion, can we have more digits shown on the LB? Say 3 or 4 like most other competitions. Thanks</p>",
      "votes": 5,
      "replies": [
        {
          "id": 417912,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-09T01:07:09.013000",
          "content": "<p>Things saturated rather quickly! I increased digits to 3, and will keep it there for most of the competition. </p>\n\n<p>This is one where I hope people rely heavily on local validation anyway. :-)</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 443853,
      "author_name": "hey!2019",
      "author_url": "",
      "post_date": "2018-12-22T15:30:49.510000",
      "content": "<p>my first kaggle competition! go go go</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 418186,
      "author_name": "NarahariBM",
      "author_url": "",
      "post_date": "2018-11-09T12:51:33.093000",
      "content": "<p>Hi Inversion,</p>\n\n<p>External Data is not allowed, fine. \nBut can you please add Wiki-text/allow us to use some standard datasets from which we can train our own language models and do transfer learning.\nWhat is the point behind not allowing language models? Isn't it a loss for the organizer when participants are handicapped by not being allowed to use state of art capability?\nNot to mention amount of learning that will be missed due to this?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 418453,
          "author_name": "IridiumBlue",
          "author_url": "",
          "post_date": "2018-11-09T23:04:59.293000",
          "content": "<p>I had the same concern earlier, as for standard datasets and language models, we do have use of all the packages and associated data that are part of Kaggle Docker.   </p>\n\n<p>See <a href=\"https://github.com/Kaggle/docker-python/blob/master/Dockerfile\">https://github.com/Kaggle/docker-python/blob/master/Dockerfile</a>, especially this line : </p>\n\n<p>python -m nltk.downloader -d /usr/share/nltk_data abc alpino averaged_perceptron_tagger \\\n    basque_grammars biocreative_ppi bllip_wsj_no_aux \\\n    book_grammars brown brown_tei cess_cat cess_esp chat80 city_database cmudict \\\n    comtrans conll2000 conll2002 conll2007 crubadan dependency_treebank \\\n    europarl_raw floresta gazetteers genesis gutenberg \\\n    ieer inaugural indian jeita kimmo knbc large_grammars lin_thesaurus mac_morpho machado \\\n    masc_tagged maxent_ne_chunker maxent_treebank_pos_tagger moses_sample movie_reviews \\\n    mte_teip5 names nps_chat omw opinion_lexicon paradigms \\\n    pil pl196x porter_test ppattach problem_reports product_reviews_1 product_reviews_2 propbank \\\n    pros_cons ptb punkt qc reuters rslp rte sample_grammars semcor senseval sentence_polarity \\\n    sentiwordnet shakespeare sinica_treebank smultron snowball_data spanish_grammars \\\n    state_union stopwords subjectivity swadesh switchboard tagsets timit toolbox treebank \\\n    twitter_samples udhr2 udhr unicode_samples universal_tagset universal_treebanks_v20 \\\nvader_lexicon verbnet webtext word2vec_sample wordnet wordnet_ic words ycoe &amp;&amp; \\</p>\n\n<p>So basically any package you might want, and if you need another - I believe the contest runners are open to suggestions to add them.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 417985,
      "author_name": "liuhao",
      "author_url": "",
      "post_date": "2018-11-09T04:36:46.333000",
      "content": "<p>Hello, this challenge is forbiden for using external data, but if pretrained model or transfer learning is possible in this challenge?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 426118,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-22T16:46:15.637000",
          "content": "<p>No, since there would be no way incorporate these within the competition constraints (i.e., if you create a kernel that uses an additional dataset, you won't be able to submit the kernel output).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 417161,
      "author_name": "Sundaresh",
      "author_url": "",
      "post_date": "2018-11-07T21:44:50.890000",
      "content": "<p>Hi - can we work on multiple kernels (as long as we make only one submission) ?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 417907,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-09T00:59:25.333000",
          "content": "<p>Hi @Sundaresh -</p>\n\n<p>Yes, you can create multiple Kernels, but at the end of the competition you will only be able to select two for final scoring.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 417173,
      "author_name": "IridiumBlue",
      "author_url": "",
      "post_date": "2018-11-07T22:11:32.567000",
      "content": "<p>What are the constrains on local data?  I get that we are limited in our access to large external datasets, but suppose we have an array of 100 magic numbers that makes our kernel a winner?  10,000?   1,000,000 ?   </p>",
      "votes": 4,
      "replies": [
        {
          "id": 417910,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-09T01:02:53.720000",
          "content": "<p>You can utilize any data that is generated during the Kernel run. Most people, for example, will generate new data as part of the feature creation process. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423712,
          "author_name": "Steve Draper",
          "author_url": "",
          "post_date": "2018-11-18T22:59:56.467000",
          "content": "<p>Code is not generated during a Kernel run.  I think that the OP is asking is 'what stops me using FAR more compute offline to pre-train a network [with arbitrary external data] and write a script that extracts all its parameters as constants in code, which I then paste into the kernel and instantiate my model weights from?</p>\n\n<p>Obviously this is intended to be disallowed, but its somewhat subjective where the boundary is.  For instance model parameters vs model hyper-parameters - I'm pretty sure most high ranked kernels will be using hyper-parameter values that result from far more extensive out-of-kernel training.</p>\n\n<p>In general it's not possible to police this sensibly, but if the rules were clear on the appropriate demarkation, it could at least make the final top-scoring kernels subject to a manual within-the-spirit test...</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 434630,
          "author_name": "Guilherme Peixoto",
          "author_url": "",
          "post_date": "2018-12-06T17:42:03.347000",
          "content": "<p>This is really tricky. eg: a list of stop words - shall one use it - would <em>easily</em> fit in memory and in code - say a list or set with 150~ or so items. Does it configure as \"unallowed use of external data\"? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 443466,
          "author_name": "eee",
          "author_url": "",
          "post_date": "2018-12-21T17:40:49.550000",
          "content": "<p>I'm worried if this still hasn't been addressed properly. There are so many people using magic number lists and misspell dictionaries for preprocessing functions, which are already predefined and copy pasted into the kernel. In theory the majority of these could be written out and thought of by anyone with any idea of preprocessing, but if it falls under external data sources then kernels will need to be rewritten. Are written lists / dictionaries considered as external? If not then is there a limit to these? Would be great to know. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 458591,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-01-20T02:07:00.683000",
          "content": "<blockquote>\n  <p>There are so many people using magic number lists and misspell dictionaries for preprocessing functions, which are already predefined and copy pasted into the kernel. </p>\n</blockquote>\n\n<p>I totally agree with you. </p>\n\n<blockquote>\n  <p>Are written lists / dictionaries considered as external? If not then is there a limit to these? Would be great to know.  </p>\n</blockquote>\n\n<p>I am worried too.  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 417143,
      "author_name": "king-menin",
      "author_url": "",
      "post_date": "2018-11-07T20:26:02.900000",
      "content": "<p>can i use packages installed from pip? where i can see list of packages that i can use?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 417905,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-09T00:56:41.133000",
          "content": "<p>You will not be able to make submissions if you <code>pip install</code> libraries into your kernel.</p>\n\n<p>You can see the list of available packages in the Kaggle Docker image here:</p>\n\n<p><a href=\"https://github.com/Kaggle/docker-python\">https://github.com/Kaggle/docker-python</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 446977,
      "author_name": "JazzStandards",
      "author_url": "",
      "post_date": "2018-12-29T00:01:18.370000",
      "content": "<p>It would be great to get an answer to this question:</p>\n\n<p><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75816#446311\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75816#446311</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 462417,
          "author_name": "IridiumBlue",
          "author_url": "",
          "post_date": "2019-01-28T09:11:15.757000",
          "content": "<p>The essential problem facing the Kaggle and Quora ppl is that 'code' and 'data' are artificial distinctions.   Code is data.   </p>\n\n<p>So the question reduces to : How big can our code be?  And more precisely, how big can our code be at maximum entropy (compressed to the max.)</p>\n\n<p>They don't want to answer this question, and I suppose I don't blame them.   If they declare it to be, say, 1 Gig, people will start using all that space to import 'data' (in the loose sense of the term.)</p>\n\n<p>If they make it too low, they might clip somebody who is using their own hand-rolled framework comprising 100,000 lines.</p>\n\n<p>So they opt for silence, so nobody can game the restriction.</p>\n\n<p>The signal that sends to us participants is  : Be careful.   A list of the five parts of speech, \"Noun, Adjective, etc.\" is obviously OK.    A list of 20 dirty words is probably OK.   A stop-word list is big - so watch out and use one already in the kernel.    </p>\n\n<p>A list of 1000 toxic verbs?   Thin ice.</p>\n\n<p>200 MB of matrix coefficients?   Red light.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433783,
      "author_name": "Ee Kin Chin",
      "author_url": "",
      "post_date": "2018-12-05T13:28:26.197000",
      "content": "<p>In order for the \"Submit to Competition\" button to be active after the Kernel commit, the following conditions must be met:</p>\n\n<p>CPU Kernel &lt;= 6 hours run-time\nGPU Kernel &lt;= 2 hours run-time\nNo internet access enabled\nNo multiple data sources enabled\nNo custom packages\nSubmission file must be named \"submission.csv\"</p>\n\n<p>Obviously there is a trick to go around some of this rules, simply copy pasting everything to the notebook and there wouldn't be a need to include data sources, and as for the run time requirements, we could always copy paste results in another kernel-&gt;commit and the \"Submit to Competition\" button will be active for you to try out your submission.csv .</p>\n\n<p>So how strict are the rules for submission before and when we reach stage two?  <a href=\"/inversion\">@inversion</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 433896,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-12-05T16:15:35.510000",
          "content": "<p>If the intent of the rules are bypassed, the submitting team risks being removed from the leaderboard.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 437212,
          "author_name": "berner",
          "author_url": "",
          "post_date": "2018-12-11T14:55:33.823000",
          "content": "<p>I get:\n<code>\nYour runtime of ... minutes exceeds the CPU kernel max of 120 minutes\n</code>\nwhy?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 455995,
          "author_name": "JM100",
          "author_url": "",
          "post_date": "2019-01-14T23:07:00.640000",
          "content": "<p>becouse rules are : CPU Kernel &lt;= 6 hours run-time, GPU Kernel &lt;= 2 hours run-time</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 426291,
      "author_name": "IridiumBlue",
      "author_url": "",
      "post_date": "2018-11-23T02:33:08.757000",
      "content": "<p>Can you give a quick summary of the stage-1/stage-2 logistics for those of us new to Kaggle?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 429460,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2018-11-28T22:18:23.753000",
          "content": "<p>Because this is a <strong>code competition</strong>, the \"two stage\" format simply means that on the final submission deadline (review the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification#Timeline\">Timeline</a> page for specifics), you must select your final submission(s) from the kernel(s) that will be run against the unseen private test set. \"Stage 1\" is effectively that final submission deadline, and \"Stage 2\" is the act of our platform running all submissions' code against the unseen private test set. This is specified in the last question on the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification#Kernels-FAQ\">Kernel FAQ</a>. The Stage 2 process will likely require 2 weeks of processing and verification time following final submission deadline.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 439075,
          "author_name": "Shahebaz Mohammad",
          "author_url": "",
          "post_date": "2018-12-14T17:32:50.123000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> We are experiencing different execution times for kernels and non-reproducable results with pytorch/keras. Will kaggle do multiple runs of kernels to average for stage 2 or just one time run. </p>\n\n<p>If its a onetime run I am wondering if CuDNN randomness could play a luck factor at top LB positions as our scores are very close. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442766,
          "author_name": "Bai",
          "author_url": "",
          "post_date": "2018-12-20T13:18:08.443000",
          "content": "<p>Is there a time limit for the second phase?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 425264,
      "author_name": "SEU.Zhou",
      "author_url": "",
      "post_date": "2018-11-21T11:09:52.363000",
      "content": "<p>The rules state that no external data could be used. I just want to know that can we use pretrained weights which needs to download thorough the Internet for transfer learning? Thx.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 426126,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-22T16:52:43.930000",
          "content": "<p>There would be no way incorporate these within the competition constraints (i.e., if you create a kernel that uses an additional dataset or via kernel internet access, you won't be able to submit the kernel output). </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 444139,
      "author_name": "Mark Peng",
      "author_url": "",
      "post_date": "2018-12-23T10:11:38.930000",
      "content": "<p>Hi <a href=\"/inversion\">@inversion</a>,</p>\n\n<p>Does the GPU kernel in this competition using the latest Kaggle GPU image?</p>\n\n<p>I made a PR to add <code>cupy</code> and <code>pynvrtc</code>, and also got merged to master branch:\n<a href=\"https://github.com/Kaggle/docker-python/blob/master/gpu.Dockerfile#L57-L58\">https://github.com/Kaggle/docker-python/blob/master/gpu.Dockerfile#L57-L58</a></p>\n\n<p>But from the kernel it says that the package is not found, any suggestions?</p>\n\n<pre><code>ModuleNotFoundError: No module named 'cupy'\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 455748,
          "author_name": "Mark Peng",
          "author_url": "",
          "post_date": "2019-01-14T13:38:21.300000",
          "content": "<p><a href=\"/inversion\">@inversion</a> Could you give some reply to this problem? It is still not working. Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 443400,
      "author_name": "redshellspy",
      "author_url": "",
      "post_date": "2018-12-21T14:57:50.663000",
      "content": "<p>Hi, inversion, \nIn the data page it is mentioned : </p>\n\n<p>This file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions.</p>\n\n<p>what does public leaderboard data remains the same for both versions mean?</p>\n\n<p>does this mean out of 376k examples in private part  56k examples will be same as the public leaderboard and 320k will be different?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 424932,
      "author_name": "Rudy Gilman",
      "author_url": "",
      "post_date": "2018-11-20T22:39:42.040000",
      "content": "<p><a href=\"/inversion\">@inversion</a> can we get an answer on the use of pretrained weights? Would like to use transfer learning approach.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 426124,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2018-11-22T16:51:11.847000",
          "content": "<p>There would be no way incorporate these within the competition constraints (i.e., if you create a kernel that uses an additional dataset, you won't be able to submit the kernel output). The exception would be tools that are contained in the Kaggle Docker packages.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 473598,
      "author_name": "eni",
      "author_url": "",
      "post_date": "2019-02-18T08:40:29.283000",
      "content": "<p>Welcoe too.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 464299,
      "author_name": "Marcin Łesek",
      "author_url": "",
      "post_date": "2019-01-31T15:01:35.580000",
      "content": "<p>My first challenge here on Kaggle, hope to help you and learn new skills. :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 462643,
      "author_name": "madara",
      "author_url": "",
      "post_date": "2019-01-28T16:29:16.690000",
      "content": "<p><a href=\"/inversion\">@inversion</a>\nHI\nIt is allowed loading \"kernel output files\" from my previous Quora kernels (my work) as input files?\nbecause it seems that this prevents me to submit output   predictions file! ( \"Submit to competition\" button remains inactive)</p>\n\n<p>thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 462418,
      "author_name": "IridiumBlue",
      "author_url": "",
      "post_date": "2019-01-28T09:13:00.013000",
      "content": "<p>Can single participants safely ignore the merger deadline?  No team, no worries?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 461642,
      "author_name": "hahaha",
      "author_url": "",
      "post_date": "2019-01-26T16:16:59.567000",
      "content": "<p><a href=\"/inversion\">@inversion</a>\nHow can I set num_workers for multiprocessing for stage2? \nwhen I tested it in kernel, cpu_core isn't equal.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 461662,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2019-01-26T16:53:37.183000",
          "content": "<p>What values are you getting?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 461758,
          "author_name": "hahaha",
          "author_url": "",
          "post_date": "2019-01-26T23:47:07.507000",
          "content": "<p>number of cpu cores, \nbut it isn't fixed (2 or 4)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 461590,
      "author_name": "Zoya",
      "author_url": "",
      "post_date": "2019-01-26T13:03:26.553000",
      "content": "<p><a href=\"/inversion\">@inversion</a> I am trying to build a model using the Google Pre Trained Word Embeddings but my kernel dies just at the end! I don't understand what's happening!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 461660,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2019-01-26T16:53:12.157000",
          "content": "<p>There have been some operational issues with kernels, but should be resolved now. Can you try again and let me know what happens? Thx.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 457529,
      "author_name": "gody7334",
      "author_url": "",
      "post_date": "2019-01-17T16:20:23.017000",
      "content": "<p>HI\nIs WordNet synsets allowed in this competition?\nAs I can run wordnet synset using Kaggle kernel, I will assume its OK, just want to double check.\nThanks,</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 456831,
      "author_name": "Endi Niu",
      "author_url": "",
      "post_date": "2019-01-16T16:04:51.360000",
      "content": "<p>When clicking 'submit to competition', it always redirect to 404 page. Could anyone help?</p>\n\n<p>Update: The reason is I saved submission file to different location like <code>sub.to_csv(\"experiment_folder/submission.csv\", index=False)</code>. I think we must do <code>sub.to_csv(\"submission.csv\", index=False)</code>, and I finally submitted successfully.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 456832,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-01-16T16:08:01.153000",
          "content": "<p><a href=\"/inversion\">@inversion</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 456849,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2019-01-16T16:46:51.847000",
          "content": "<p>What happens if you create a brand new kernel and just read/submit the sample submission?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 457125,
          "author_name": "Endi Niu",
          "author_url": "",
          "post_date": "2019-01-17T01:31:33.397000",
          "content": "<p>Thanks for reply. I succeeded to submit sample submission.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 453076,
      "author_name": "aintnosunshine",
      "author_url": "",
      "post_date": "2019-01-09T16:26:03.930000",
      "content": "<p>Is data augmentation ( by synonyms , from embedding ) is allowed in this competition or will it be considered as external data ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 451935,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-07T23:24:18.410000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 451932,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-07T23:17:43.210000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 451357,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-07T00:03:34.540000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450254,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-04T15:15:22.533000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 449977,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-04T03:59:45.917000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446681,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-28T13:39:35.507000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 446019,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-27T11:01:58.260000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 444396,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-24T00:11:58.493000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 442557,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-20T05:51:16.840000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 441155,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-18T10:34:40.537000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 443858,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-22T15:50:43.117000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 440682,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-17T22:08:42.677000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 440830,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-18T02:21:16.650000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 441631,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-18T21:38:54.340000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 440329,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-17T12:09:00.173000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 442503,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-20T03:11:27.540000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 442777,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-20T13:40:39.937000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 435382,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T23:44:35.633000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 444861,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-25T02:15:54.603000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 434910,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-07T06:22:45.743000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 431921,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-03T05:15:52.257000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 444873,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-25T03:00:38.500000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 429367,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-28T18:43:50.367000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 427359,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-25T10:20:30.683000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 426826,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-24T00:17:44.563000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 450636,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-05T12:42:45.573000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 426232,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-22T23:46:47.633000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422090,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-15T19:07:57.437000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 419442,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-12T01:04:35.627000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 417514,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-08T12:01:26.420000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 417913,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-09T01:07:56.013000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 450320,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-04T17:18:27.700000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 450305,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-04T16:52:53.470000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 418573,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-10T06:52:47.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 418023,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-09T05:51:20.217000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 417168,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-07T21:58:01.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 473458,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-18T03:08:32.027000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "416485": "Welcome to the Quora Insincere Questions Classification Challenge!\n\nIn this challenge, you are tasked with developing models that identify insincere questions submitted to Quora.\n\nThis is a Kernels-Only competition - you should familiarize yourself with the [Kernels FAQ](https://www.kaggle.com/c/quora-insincere-questions-classification#Kernels-FAQ) for important information about Kernels runtime and environment constraints to be able to make a submission.\n\nYou'll also want to read through the [Data](https://www.kaggle.com/c/quora-insincere-questions-classification/data) page to understand the target variable, supplementary data that's provided, etc.\n\nPlease feel free to ask your questions in this thread.\n\nGood luck!",
    "417390": "Looking at the training data, am I the only one who sees lots of entries which seem to have been tagged w/o careful review?  Examples - \n\n0098e5b8335a62791abf,Is Dr. Steven Greer an extremely good fraudster with his claims of ufo' s and aliens?,1\n009afb0a3716d819bf7e,What did Trump and Sessions possibly gain from firing McCabe with less than 24 hours until retiring?,1\nWhy do progressive activists feel that tearing down antique historic but politically incorrect statues accomplishes a positive good?,1\ndc002c8cff10ddc797dc,Does Trump's inaction on his own justice department's indictment of Russian cyber attacks mean he has violated his oath of office?,1\ndc16793a4d38a9a27803,How do black people celebrate Martin Luther King Jr's birthday?,1\n00b92227d01a117d9c85,\"Do white people in America think they are privileged? Is that a real thing going on, or just a bad meme?\",1\ndbd98d3003044f51791a,Do Democrats have a plan to curb violence in Chicago?,1\n00c6da623540ba7d0a9d,\"What would have to be proven to show that Cohen is guilty of influence peddling, a crime? Any more than is already a matter of public record?\",1\ndc44f81813106a237458,Why doesn't China apologize for invading Tibet? They did it even before the Sino-Indian War.,1\n00c7fbcb0495b64ed7c2,\"If Donald Trump pardons himself, would that count as an admission of guilt and make it easier for the state attorney general to prosecute on related but not identical charges?\",1\n00cc8173e3c8a2a5c3fd,Can asperger's recognise sarcasm in relationships?,1 \n00df0e17c0de8d326227Do some atheists feel guilty about indoctrinating their children with atheism and not giving them hope and a moral compass?,1\ndcb2c4b6be697b16c9e3,How do I avoid losing the argument when I try to defend ethnic groups with statistically proven higher criminality rates? Am I wrong to defend these ethnic groups in the face of hard facts?,1\ndcbaf4d21e0b29592e8c,Does the conservatives’ view on climate change contradict their core conservative values?,1\ndc972c2e4f084dc8fdb8,\"Why do so many Americans get puzzled when I ask for a serviette in their restaurants, but know right away what I am talking about when I ask for a Napkin (aka diaper)?\",1\n00df34daff54eef050f7What is it like to be a gay in IIT Madras?,1\n\n\nBut here we have  : dbae47ac9a7adf7f178cWhat is it like to be gay in Southern California?,0\n\n(So - gay as an adjective is OK, but not as a noun?)\n\nThis is about 10 % of the 1's, near as I can reckon.    Yes, these are ham-fisted questions about delicate topics.   Yes, they could be worded more gently.   But they do not harass nor disparage, they certainly appear to be in good faith and seeking real answers.\n\nThere's a real garbage-in garbage-out problem here.   The dataset appears not to have been human-curated at all, but rather from an algorithm which is surfacing false positives.   Or from a curator who is all nerve endings about certain issues.\n\nI was very excited to enter (and win!!!) this competition but I am stopped dead in my tracks by bad training data.  My complaint isn't about politics, I could not describe to a human with a PhD in English Literature how to arrive at the conclusions given in this sample.\n\nTrying to code it is a fool's errand, to put it mildly.\n\nNot convinced?   How about this one - \n\ndcbfe10facdd66c91eca,Why aren't there any women in the Navy SEALS?,1\n\nExplain that one.   To the US Navy.    https://www.navytimes.com/news/your-navy/2018/02/16/two-women-could-enter-navy-special-operations-training-this-year/\n\n\n\n\n\n\n",
    "422736": "Seems most of us use CUDNN package which has strong randomness for each run (but very fast). Is that possible in stage 2, you run the model a few times and take an average?",
    "416988": "Great to have another NLP-related competition on Kaggle! Surely going to give this a go!",
    "416747": "Can the docker image change during the course of the competition? E.g. can it happen that some interesting library will not be available at the beginning, but and made available later? If so, do I have to monitor commits to the official docker image to make myself aware of this?\n",
    "416527": "In stage 2 of the competition you rerun the kernels on a test set of approx 376k vs approx 56k in the first stage. Is there any adjustment made to the available run times in the kernels, or do they still have to run in 6 hours (or 2 hours for GPU)?",
    "424161": "@inversion Are vectors and other models from spaCy fair game??\n\nOther people have asked same thing here: https://www.kaggle.com/jpmiller/bonus-vectors-with-spacy/comments\n\nThanks!",
    "417221": "Hi inversion, can we have more digits shown on the LB? Say 3 or 4 like most other competitions. Thanks",
    "443853": "my first kaggle competition! go go go",
    "418186": "Hi Inversion,\n\nExternal Data is not allowed, fine. \nBut can you please add Wiki-text/allow us to use some standard datasets from which we can train our own language models and do transfer learning.\nWhat is the point behind not allowing language models? Isn't it a loss for the organizer when participants are handicapped by not being allowed to use state of art capability?\nNot to mention amount of learning that will be missed due to this?",
    "417985": "Hello, this challenge is forbiden for using external data, but if pretrained model or transfer learning is possible in this challenge?",
    "417161": "Hi - can we work on multiple kernels (as long as we make only one submission) ?\n",
    "417173": "What are the constrains on local data?  I get that we are limited in our access to large external datasets, but suppose we have an array of 100 magic numbers that makes our kernel a winner?  10,000?   1,000,000 ?   ",
    "417143": "can i use packages installed from pip? where i can see list of packages that i can use?",
    "446977": "It would be great to get an answer to this question:\n\nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/75816#446311",
    "433783": "In order for the \"Submit to Competition\" button to be active after the Kernel commit, the following conditions must be met:\n\nCPU Kernel &lt;= 6 hours run-time\nGPU Kernel &lt;= 2 hours run-time\nNo internet access enabled\nNo multiple data sources enabled\nNo custom packages\nSubmission file must be named \"submission.csv\"\n\nObviously there is a trick to go around some of this rules, simply copy pasting everything to the notebook and there wouldn't be a need to include data sources, and as for the run time requirements, we could always copy paste results in another kernel-&gt;commit and the \"Submit to Competition\" button will be active for you to try out your submission.csv .\n\nSo how strict are the rules for submission before and when we reach stage two?  @inversion",
    "426291": "Can you give a quick summary of the stage-1/stage-2 logistics for those of us new to Kaggle?",
    "425264": "The rules state that no external data could be used. I just want to know that can we use pretrained weights which needs to download thorough the Internet for transfer learning? Thx.",
    "444139": "Hi @inversion,\n\nDoes the GPU kernel in this competition using the latest Kaggle GPU image?\n\nI made a PR to add `cupy` and `pynvrtc`, and also got merged to master branch:\nhttps://github.com/Kaggle/docker-python/blob/master/gpu.Dockerfile#L57-L58\n\nBut from the kernel it says that the package is not found, any suggestions?\n\n    ModuleNotFoundError: No module named 'cupy'\n    ",
    "443400": "Hi, inversion, \nIn the data page it is mentioned : \n\nThis file will have ~56k rows in stage 1 and ~376k rows in stage 2. The public leaderboard data remains the same for both versions.\n\nwhat does public leaderboard data remains the same for both versions mean?\n\ndoes this mean out of 376k examples in private part  56k examples will be same as the public leaderboard and 320k will be different?",
    "424932": "@inversion can we get an answer on the use of pretrained weights? Would like to use transfer learning approach.",
    "473598": "Welcoe too.",
    "464299": "My first challenge here on Kaggle, hope to help you and learn new skills. :)",
    "462643": "@inversion\nHI\nIt is allowed loading \"kernel output files\" from my previous Quora kernels (my work) as input files?\nbecause it seems that this prevents me to submit output   predictions file! ( \"Submit to competition\" button remains inactive)\n\nthanks",
    "462418": "Can single participants safely ignore the merger deadline?  No team, no worries?",
    "461642": "@inversion\nHow can I set num_workers for multiprocessing for stage2? \nwhen I tested it in kernel, cpu_core isn't equal.",
    "461590": "@inversion I am trying to build a model using the Google Pre Trained Word Embeddings but my kernel dies just at the end! I don't understand what's happening!",
    "457529": "HI\nIs WordNet synsets allowed in this competition?\nAs I can run wordnet synset using Kaggle kernel, I will assume its OK, just want to double check.\nThanks,",
    "456831": "When clicking 'submit to competition', it always redirect to 404 page. Could anyone help?\n\nUpdate: The reason is I saved submission file to different location like `sub.to_csv(\"experiment_folder/submission.csv\", index=False)`. I think we must do `sub.to_csv(\"submission.csv\", index=False)`, and I finally submitted successfully.",
    "453076": "Is data augmentation ( by synonyms , from embedding ) is allowed in this competition or will it be considered as external data ?",
    "451935": "Could you please install GPU version of Apache MXNet? Currently, only CPU version is installed, and if I try to use GPU I get an error. Here is  the example:\n```\nimport mxnet as mx\na = mx.nd.array([1,2,34], mx.gpu(0))\n```\nThe error message is:\n\n```\nMXNetError: [23:21:15] src/storage/storage.cc:137: Compile with USE_CUDA=1 to enable GPU usage Stack trace returned 10 entries: [bt] (0) /opt/conda/lib/python3.6/site-packages/mxnet/libmxnet.so(+0x1e945a) [0x7f45bb45445a] [bt] (1) /opt/conda/lib/python3.6/site-packages/mxnet/libmxnet.so(+0x1e9ac1) [0x7f45bb454ac1] [bt] (2) /opt/conda/lib/python3.6/site-packages/mxnet/libmxnet.so(+0x2f83205) [0x7f45be1ee205]\n```\n\nTo do so, MXNet should be installed with `pip install mxnet-cu90mkl` (for CUDA 9.0 and MKL support).\nThank you.",
    "451932": "Could you please allow usage of http://gluon-nlp.mxnet.io/ package? It is state of the art NLP library for Apache MXNet framework, and it is as easy to install as `pip install gluonnlp`. Right now I can import it for CPU instance only.",
    "451357": "@inversion the training set looks like it contains quite a few false negatives (which is pretty understandable!). Can you tell us if the test set labelled in the same way or has it been human validated?\nThanks for setting this up.",
    "450254": "@inversion is it fair to include pseudo labeling manually ? i.e. print the prediction then add it again to the training using copy paste ?",
    "449977": "HI can someone help me with an error. Its weird that I can't post my problems in discussions as well. When I try to submit my output, it says that I cannot use docker image. What does it mean?\n\n",
    "446681": "i am a new nlper! quora  competition is very funny!",
    "446019": "@inversion \nI find an interesting thing, the distribution between training data and testing data is different.\nSo I want to know whether the public and private testing data are derived from a large test set, then randomly cut into two parts, or derived from two data sets? ",
    "444396": "Hey, This is Chandrashekhar Adhage...Excited for this Competition..:)",
    "442557": "@inversion Are the questions in languages other than english correctly labeled (insecure or not)? When labeling the training data, were different models used for different languages? ",
    "441155": "Is it fair to manually remove or remark false positives of the dataset inside the code of the kernel?\nfor example:\n```\ndata[data['question_text'] in [TEXTS TO MANUALLY MARK]]['target']=0\n```",
    "440682": "Hi, this is my first Kaggle competition. I submitted and got a score of 0.37. Does this mean I got a 37% classification accuracy?",
    "440329": "In the Private LB phase, will our model be re-run? Or use the parameters of the current optimal model to directly predict the results of the new test data?",
    "435382": "Considering the fact, that the Kernel FAQ states \"No multiple data sources enabled\": Does this apply to the given word vector files? Meaning, are we only allowed to use one of those word vector files? \n\nAnd if I create a dict object in the kernel which contains words to be replaced, does this count as a data source? This question pertains to any form of data manipulation of the training/test data which is based on a manual analysis of this data and results in some form of data collection.\n\nAnd if such analysis data - like the dict above - counts as a data source, does this lead to a conflict with loaded word vector files - meaning only the dict or the word vectors are allowed?",
    "434910": "can i train different models on different kernels, and combine these pretrained models generated on final kernel?",
    "431921": "I recently started working on ML problems as became more aware of ML.Net (Microsoft's open source ML library). Is there anyway i can get score for my submission which is in Output folder but was uploaded from local drive since I had worked on this locally. I understand that I did not adhere to rules, and I am fine with being disqualified. But I would love to see what was the score for my predictions.",
    "429367": "Hi. GloVe: Global Vectors for Word Representation using nltk.tokenize.stanford.StanfordTokenizer, but we can't use it without stanford-postagger.jar (3,49 Mb) and os.environ['JAVAHOME'] = java_path. Can you add this file to data?\nZip with file: https://nlp.stanford.edu/software/stanford-postagger-2018-10-16.zip",
    "427359": "Could you please update the python package like \"fastai\",  thanks :)",
    "426826": "@inversion I am curious to know how exactly the overall F1 score is being calculated here? Is it micro, macro, or some other?",
    "426232": "Insincere questions are unsafe.",
    "422090": "Can I use the information on the test dataset as part of my training? Like to extract all the words from it to the model's vocabulary. Thanks.",
    "419442": "Intresting NLP-related competition on Kaggle! \n\n",
    "417514": "Am i eligible to use embeddings which trained on competition data?",
    "450320": "@inversion I have a question that worries me a bit:\nAbove private LB it says \"The private leaderboard is calculated over the same rows as the public leaderboard in this competition.\"\n\nIt is now clear that kernels will be re-run in phase 2 to predict labels on all 376k rows.\nBut then to get f1 score there are two options:\n(1) 376k rows are partitioned into 56k and 320k to get f1 score of public and private LB respectively.\n(2) 376k row are reorganized into 56k and 376k to get f1 score of public and private LB respectively. \n\nWhich of (1) or (2) will be the procedure?",
    "450305": "In description it says test data has ~56k rows in stage 1 and ~376k rows in stage 2. \nThus, ~320k extra rows will be included for stage 2.\n\nAre these 56k and 320k rows random samples of the full test dataset of 376k rows?\n... or are they two separate datasets, possibly with a very different frequency of class 1 (insincere questions) ?",
    "418573": "",
    "418023": "",
    "417168": "",
    "473458": "Thanks a lot"
  }
}