{
  "id": 31368,
  "title": "External Data?",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/31368",
  "author_name": "",
  "post_date": "2017-04-09T18:48:47.394365200Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>It isn't super clear to me based on the rules whether external data is allowed. Can someone clarify for me? Thank you!</p>",
  "messages": [
    {
      "id": "173988",
      "postDate": "04/09/2017 18:48:47",
      "content": "<p>It isn't super clear to me based on the rules whether external data is allowed. Can someone clarify for me? Thank you!</p>",
      "rawMarkdown": "It isn't super clear to me based on the rules whether external data is allowed. Can someone clarify for me? Thank you!",
      "votes": null
    },
    {
      "id": "175672",
      "postDate": "04/16/2017 21:46:55",
      "content": "<p>Any updates regarding this?</p>",
      "rawMarkdown": "Any updates regarding this?",
      "votes": null
    },
    {
      "id": "179043",
      "postDate": "04/29/2017 21:40:05",
      "content": "<p>General rule:</p>\n\n<blockquote>\n  <p>Unless otherwise expressly stated on the Competition Website,\n  Participants must not use data other than the Data to develop and test\n  their models and Submissions.</p>\n</blockquote>",
      "rawMarkdown": "General rule:\n\n> Unless otherwise expressly stated on the Competition Website,\n> Participants must not use data other than the Data to develop and test\n> their models and Submissions.",
      "votes": null
    },
    {
      "id": "179068",
      "postDate": "04/30/2017 02:06:38",
      "content": "<p>External data is not allowed. But external models are... as long as they're shared?</p>\n\n<p>That doesn't make sense to me, as a general Kaggle rule. After all, what are external models except curated, aggregated, and summarized external data?</p>\n\n<p>By the current rules, one would get disqualified by means of the no external data clause if they trained their model on a freely available dataset of 100000 sea lions. However, if they trained a <em>portion</em> of their model on the <strong>same exact dataset</strong>, then open-sourced <em>that portion</em> of their model as an \"external model\" (minus any fine tuning) 1 week before the competition ending, that would be completely acceptable. What gives?</p>\n\n<p>For the record, I'm a proponent for allowing <em>both</em> external models as well as external data, so long as they're shared.</p>",
      "rawMarkdown": "External data is not allowed. But external models are... as long as they're shared?\n\nThat doesn't make sense to me, as a general Kaggle rule. After all, what are external models except curated, aggregated, and summarized external data?\n\nBy the current rules, one would get disqualified by means of the no external data clause if they trained their model on a freely available dataset of 100000 sea lions. However, if they trained a *portion* of their model on the **same exact dataset**, then open-sourced *that portion* of their model as an \"external model\" (minus any fine tuning) 1 week before the competition ending, that would be completely acceptable. What gives?\n\nFor the record, I'm a proponent for allowing *both* external models as well as external data, so long as they're shared.",
      "votes": null
    },
    {
      "id": "179070",
      "postDate": "04/30/2017 02:18:10",
      "content": "<p>Interesting loophole - perhaps the rules should be clarified that only 'well known' models should be allowed?  </p>\n\n<p>I guess a freshly trained model to be applied (mostly) for a Kaggle competition wouldn't be a *pre*trained model per se...</p>",
      "rawMarkdown": "Interesting loophole - perhaps the rules should be clarified that only 'well known' models should be allowed?  \n\nI guess a freshly trained model to be applied (mostly) for a Kaggle competition wouldn't be a *pre*trained model per se...",
      "votes": null
    },
    {
      "id": "179103",
      "postDate": "04/30/2017 05:47:46",
      "content": "<blockquote>\n  <p>Perhaps the rules should be clarified that only 'well known' models should be allowed?</p>\n</blockquote>\n\n<p>I dunno, that seems counter-intuitive. If I were a project-owner cooperating with (paying?) Kaggle, I'd want to ensure whatever I get out of the deal was something that would be beneficial to me as a corporation... and not just a fun + entertaining project for Kagglers. In other words, if <strong>freely</strong> available data exists, then it <strong>should</strong> be permissible to use that in competitions, the same way models are. It's up to the competitors to google for data that will help them.</p>\n\n<p>The only advantage to limiting sharing to one week ahead of deadline is to foster team-work between Kagglers so that each team competes for higher scores using each other's techniques. But if people refuse to share, or what worse--if people are further handicapped to a point where they can't even use beneficial external data, that only inhibits their ability to produce state of the art solutions... which is counter productive to the goal of the project-owners.</p>",
      "rawMarkdown": "> Perhaps the rules should be clarified that only 'well known' models should be allowed?\n\nI dunno, that seems counter-intuitive. If I were a project-owner cooperating with (paying?) Kaggle, I'd want to ensure whatever I get out of the deal was something that would be beneficial to me as a corporation... and not just a fun + entertaining project for Kagglers. In other words, if **freely** available data exists, then it **should** be permissible to use that in competitions, the same way models are. It's up to the competitors to google for data that will help them.\n\nThe only advantage to limiting sharing to one week ahead of deadline is to foster team-work between Kagglers so that each team competes for higher scores using each other's techniques. But if people refuse to share, or what worse--if people are further handicapped to a point where they can't even use beneficial external data, that only inhibits their ability to produce state of the art solutions... which is counter productive to the goal of the project-owners.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 175672,
      "author_name": "skrish13",
      "author_url": "",
      "post_date": "04/16/2017 21:46:55",
      "content": "<p>Any updates regarding this?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 179043,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "04/29/2017 21:40:05",
      "content": "<p>General rule:</p>\n\n<blockquote>\n  <p>Unless otherwise expressly stated on the Competition Website,\n  Participants must not use data other than the Data to develop and test\n  their models and Submissions.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 179068,
      "author_name": "authman",
      "author_url": "",
      "post_date": "04/30/2017 02:06:38",
      "content": "<p>External data is not allowed. But external models are... as long as they're shared?</p>\n\n<p>That doesn't make sense to me, as a general Kaggle rule. After all, what are external models except curated, aggregated, and summarized external data?</p>\n\n<p>By the current rules, one would get disqualified by means of the no external data clause if they trained their model on a freely available dataset of 100000 sea lions. However, if they trained a <em>portion</em> of their model on the <strong>same exact dataset</strong>, then open-sourced <em>that portion</em> of their model as an \"external model\" (minus any fine tuning) 1 week before the competition ending, that would be completely acceptable. What gives?</p>\n\n<p>For the record, I'm a proponent for allowing <em>both</em> external models as well as external data, so long as they're shared.</p>",
      "votes": null,
      "replies": [
        {
          "id": 179070,
          "author_name": "happycube",
          "author_url": "",
          "post_date": "04/30/2017 02:18:10",
          "content": "<p>Interesting loophole - perhaps the rules should be clarified that only 'well known' models should be allowed?  </p>\n\n<p>I guess a freshly trained model to be applied (mostly) for a Kaggle competition wouldn't be a *pre*trained model per se...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 179103,
          "author_name": "authman",
          "author_url": "",
          "post_date": "04/30/2017 05:47:46",
          "content": "<blockquote>\n  <p>Perhaps the rules should be clarified that only 'well known' models should be allowed?</p>\n</blockquote>\n\n<p>I dunno, that seems counter-intuitive. If I were a project-owner cooperating with (paying?) Kaggle, I'd want to ensure whatever I get out of the deal was something that would be beneficial to me as a corporation... and not just a fun + entertaining project for Kagglers. In other words, if <strong>freely</strong> available data exists, then it <strong>should</strong> be permissible to use that in competitions, the same way models are. It's up to the competitors to google for data that will help them.</p>\n\n<p>The only advantage to limiting sharing to one week ahead of deadline is to foster team-work between Kagglers so that each team competes for higher scores using each other's techniques. But if people refuse to share, or what worse--if people are further handicapped to a point where they can't even use beneficial external data, that only inhibits their ability to produce state of the art solutions... which is counter productive to the goal of the project-owners.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "173988": "It isn't super clear to me based on the rules whether external data is allowed. Can someone clarify for me? Thank you!",
    "175672": "Any updates regarding this?",
    "179043": "General rule:\n\n> Unless otherwise expressly stated on the Competition Website,\n> Participants must not use data other than the Data to develop and test\n> their models and Submissions.",
    "179068": "External data is not allowed. But external models are... as long as they're shared?\n\nThat doesn't make sense to me, as a general Kaggle rule. After all, what are external models except curated, aggregated, and summarized external data?\n\nBy the current rules, one would get disqualified by means of the no external data clause if they trained their model on a freely available dataset of 100000 sea lions. However, if they trained a *portion* of their model on the **same exact dataset**, then open-sourced *that portion* of their model as an \"external model\" (minus any fine tuning) 1 week before the competition ending, that would be completely acceptable. What gives?\n\nFor the record, I'm a proponent for allowing *both* external models as well as external data, so long as they're shared.",
    "179070": "Interesting loophole - perhaps the rules should be clarified that only 'well known' models should be allowed?  \n\nI guess a freshly trained model to be applied (mostly) for a Kaggle competition wouldn't be a *pre*trained model per se...",
    "179103": "> Perhaps the rules should be clarified that only 'well known' models should be allowed?\n\nI dunno, that seems counter-intuitive. If I were a project-owner cooperating with (paying?) Kaggle, I'd want to ensure whatever I get out of the deal was something that would be beneficial to me as a corporation... and not just a fun + entertaining project for Kagglers. In other words, if **freely** available data exists, then it **should** be permissible to use that in competitions, the same way models are. It's up to the competitors to google for data that will help them.\n\nThe only advantage to limiting sharing to one week ahead of deadline is to foster team-work between Kagglers so that each team competes for higher scores using each other's techniques. But if people refuse to share, or what worse--if people are further handicapped to a point where they can't even use beneficial external data, that only inhibits their ability to produce state of the art solutions... which is counter productive to the goal of the project-owners."
  },
  "source": "meta"
}