{
  "id": 70978,
  "title": "When does something become an \"external data source\"?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/70978",
  "author_name": "",
  "post_date": "2018-11-09T00:29:42.173741600Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I've noticed from some of the EDAs that have been posted so far that US politicians such as Donald Trump tend to be a common topic in questions. I thought it might be useful to have a list of US politicians from something like Open States or WikiData and a feature of whether or not the question contains a US politician's name. Obviously external data sources are not allowed though.</p>\n\n<p>But I could theoretically just make a list of US Presidential candidates and insert it in the code and that <strong>probably</strong> wouldn't be considered an \"external data source\". Would that violate the rules? Should the rules specify that wouldn't be allowed?</p>",
  "messages": [
    {
      "id": "417896",
      "postDate": "11/09/2018 00:29:42",
      "content": "<p>I've noticed from some of the EDAs that have been posted so far that US politicians such as Donald Trump tend to be a common topic in questions. I thought it might be useful to have a list of US politicians from something like Open States or WikiData and a feature of whether or not the question contains a US politician's name. Obviously external data sources are not allowed though.</p>\n\n<p>But I could theoretically just make a list of US Presidential candidates and insert it in the code and that <strong>probably</strong> wouldn't be considered an \"external data source\". Would that violate the rules? Should the rules specify that wouldn't be allowed?</p>",
      "rawMarkdown": "I've noticed from some of the EDAs that have been posted so far that US politicians such as Donald Trump tend to be a common topic in questions. I thought it might be useful to have a list of US politicians from something like Open States or WikiData and a feature of whether or not the question contains a US politician's name. Obviously external data sources are not allowed though.\n\nBut I could theoretically just make a list of US Presidential candidates and insert it in the code and that **probably** wouldn't be considered an \"external data source\". Would that violate the rules? Should the rules specify that wouldn't be allowed?",
      "votes": null
    },
    {
      "id": "418095",
      "postDate": "11/09/2018 09:26:45",
      "content": "<p>I am also wondering the same. Is e.g., using some pre-compiled vocabularies for sentiment analysis in python packages considered as external data? Or things like named entity recognition from nltk or spacy?</p>",
      "rawMarkdown": "I am also wondering the same. Is e.g., using some pre-compiled vocabularies for sentiment analysis in python packages considered as external data? Or things like named entity recognition from nltk or spacy?",
      "votes": null
    },
    {
      "id": "418354",
      "postDate": "11/09/2018 18:02:28",
      "content": "<p>In general, if it common knowledge, it would not be considered external data. </p>",
      "rawMarkdown": "In general, if it common knowledge, it would not be considered external data.",
      "votes": null
    },
    {
      "id": "418366",
      "postDate": "11/09/2018 18:27:13",
      "content": "<p>Thank you for the clarification. I guess that is a rule of thumb also for other competitions</p>",
      "rawMarkdown": "Thank you for the clarification. I guess that is a rule of thumb also for other competitions",
      "votes": null
    },
    {
      "id": "418407",
      "postDate": "11/09/2018 21:32:48",
      "content": "<p>Interesting. Would that mean a <a href=\"https://docs.openstates.org/en/latest/data/legacy-csv.html#legislators-csv\">CSV</a> of legislator names downloaded from somewhere like Open States would be allowed? </p>",
      "rawMarkdown": "Interesting. Would that mean a [CSV](https://docs.openstates.org/en/latest/data/legacy-csv.html#legislators-csv) of legislator names downloaded from somewhere like Open States would be allowed?",
      "votes": null
    },
    {
      "id": "418586",
      "postDate": "11/10/2018 07:16:00",
      "content": "<p>spacy has pretrained embeddings and neural networks that definitely should count as using exterior data. not sure about nltk</p>",
      "rawMarkdown": "spacy has pretrained embeddings and neural networks that definitely should count as using exterior data. not sure about nltk",
      "votes": null
    },
    {
      "id": "418811",
      "postDate": "11/10/2018 16:45:59",
      "content": "<p>NLTK usually resorts to some built-in vocabularies and models.</p>",
      "rawMarkdown": "NLTK usually resorts to some built-in vocabularies and models.",
      "votes": null
    },
    {
      "id": "419585",
      "postDate": "11/12/2018 08:56:05",
      "content": "<p>Can you refine this? It is hard to judge where the line is for this challenge regarding external data.</p>",
      "rawMarkdown": "Can you refine this? It is hard to judge where the line is for this challenge regarding external data.",
      "votes": null
    },
    {
      "id": "419941",
      "postDate": "11/12/2018 19:48:11",
      "content": "<p>It's difficult to provide an exact line. In general, if the information might reasonably be known by someone (e.g., U.S. state capitals), it's fair game. If it's highly unlikely the information could be known by someone without looking it up (e.g., the <em>lat/lon and average yearly temperature</em> of U.S. state capitals), then it is external data.</p>",
      "rawMarkdown": "It's difficult to provide an exact line. In general, if the information might reasonably be known by someone (e.g., U.S. state capitals), it's fair game. If it's highly unlikely the information could be known by someone without looking it up (e.g., the *lat/lon and average yearly temperature* of U.S. state capitals), then it is external data.",
      "votes": null
    },
    {
      "id": "423629",
      "postDate": "11/18/2018 18:42:31",
      "content": "<p>As this discussion also popped up somewhere else: Is it OK to use pre-trained models from e.g. NLTK or Spacy?</p>",
      "rawMarkdown": "As this discussion also popped up somewhere else: Is it OK to use pre-trained models from e.g. NLTK or Spacy?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 418095,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "11/09/2018 09:26:45",
      "content": "<p>I am also wondering the same. Is e.g., using some pre-compiled vocabularies for sentiment analysis in python packages considered as external data? Or things like named entity recognition from nltk or spacy?</p>",
      "votes": null,
      "replies": [
        {
          "id": 418586,
          "author_name": "julesgm",
          "author_url": "",
          "post_date": "11/10/2018 07:16:00",
          "content": "<p>spacy has pretrained embeddings and neural networks that definitely should count as using exterior data. not sure about nltk</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418811,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/10/2018 16:45:59",
          "content": "<p>NLTK usually resorts to some built-in vocabularies and models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418354,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "11/09/2018 18:02:28",
      "content": "<p>In general, if it common knowledge, it would not be considered external data. </p>",
      "votes": null,
      "replies": [
        {
          "id": 418366,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "11/09/2018 18:27:13",
          "content": "<p>Thank you for the clarification. I guess that is a rule of thumb also for other competitions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418407,
          "author_name": "robroseknows",
          "author_url": "",
          "post_date": "11/09/2018 21:32:48",
          "content": "<p>Interesting. Would that mean a <a href=\"https://docs.openstates.org/en/latest/data/legacy-csv.html#legislators-csv\">CSV</a> of legislator names downloaded from somewhere like Open States would be allowed? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419585,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/12/2018 08:56:05",
          "content": "<p>Can you refine this? It is hard to judge where the line is for this challenge regarding external data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419941,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "11/12/2018 19:48:11",
          "content": "<p>It's difficult to provide an exact line. In general, if the information might reasonably be known by someone (e.g., U.S. state capitals), it's fair game. If it's highly unlikely the information could be known by someone without looking it up (e.g., the <em>lat/lon and average yearly temperature</em> of U.S. state capitals), then it is external data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 423629,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/18/2018 18:42:31",
          "content": "<p>As this discussion also popped up somewhere else: Is it OK to use pre-trained models from e.g. NLTK or Spacy?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "417896": "I've noticed from some of the EDAs that have been posted so far that US politicians such as Donald Trump tend to be a common topic in questions. I thought it might be useful to have a list of US politicians from something like Open States or WikiData and a feature of whether or not the question contains a US politician's name. Obviously external data sources are not allowed though.\n\nBut I could theoretically just make a list of US Presidential candidates and insert it in the code and that **probably** wouldn't be considered an \"external data source\". Would that violate the rules? Should the rules specify that wouldn't be allowed?",
    "418095": "I am also wondering the same. Is e.g., using some pre-compiled vocabularies for sentiment analysis in python packages considered as external data? Or things like named entity recognition from nltk or spacy?",
    "418354": "In general, if it common knowledge, it would not be considered external data.",
    "418366": "Thank you for the clarification. I guess that is a rule of thumb also for other competitions",
    "418407": "Interesting. Would that mean a [CSV](https://docs.openstates.org/en/latest/data/legacy-csv.html#legislators-csv) of legislator names downloaded from somewhere like Open States would be allowed?",
    "418586": "spacy has pretrained embeddings and neural networks that definitely should count as using exterior data. not sure about nltk",
    "418811": "NLTK usually resorts to some built-in vocabularies and models.",
    "419585": "Can you refine this? It is hard to judge where the line is for this challenge regarding external data.",
    "419941": "It's difficult to provide an exact line. In general, if the information might reasonably be known by someone (e.g., U.S. state capitals), it's fair game. If it's highly unlikely the information could be known by someone without looking it up (e.g., the *lat/lon and average yearly temperature* of U.S. state capitals), then it is external data.",
    "423629": "As this discussion also popped up somewhere else: Is it OK to use pre-trained models from e.g. NLTK or Spacy?"
  },
  "source": "meta"
}