{
  "id": 204680,
  "title": "Why are competition data formats often so complicated?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/204680",
  "author_name": "",
  "post_date": "2020-12-16T10:34:20.753043300Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Might be a little off-topic, but after facing the same problem in different competitions over time, I've started wondering why the data formats in kaggle competitions are often so (unnecessarily?) complicated. <a href=\"https://www.kaggle.com/friedchips/fully-correct-hubmap-rle-encoding-and-decoding\" target=\"_blank\">These two functions</a> cost me a full evening…<br>\nLike in this competition: TIFFs in different formats, submissions as RLE strings. Why? Compressed numpy arrays (npz) would work just fine in both cases.<br>\nA lot of time and effort is spent on just handling/formatting the data correctly. I get it that this is an important aspect (maybe even <em>the</em> most important aspect) of data analysis in the real world, but here, it distracts from the real challenge, model building.<br>\nI'd be interested in your thoughts on this.</p>",
  "messages": [
    {
      "id": "1115507",
      "postDate": "12/16/2020 10:34:20",
      "content": "<p>Might be a little off-topic, but after facing the same problem in different competitions over time, I've started wondering why the data formats in kaggle competitions are often so (unnecessarily?) complicated. <a href=\"https://www.kaggle.com/friedchips/fully-correct-hubmap-rle-encoding-and-decoding\" target=\"_blank\">These two functions</a> cost me a full evening…<br>\nLike in this competition: TIFFs in different formats, submissions as RLE strings. Why? Compressed numpy arrays (npz) would work just fine in both cases.<br>\nA lot of time and effort is spent on just handling/formatting the data correctly. I get it that this is an important aspect (maybe even <em>the</em> most important aspect) of data analysis in the real world, but here, it distracts from the real challenge, model building.<br>\nI'd be interested in your thoughts on this.</p>",
      "rawMarkdown": "Might be a little off-topic, but after facing the same problem in different competitions over time, I've started wondering why the data formats in kaggle competitions are often so (unnecessarily?) complicated. [These two functions](https://www.kaggle.com/friedchips/fully-correct-hubmap-rle-encoding-and-decoding) cost me a full evening...\nLike in this competition: TIFFs in different formats, submissions as RLE strings. Why? Compressed numpy arrays (npz) would work just fine in both cases.\nA lot of time and effort is spent on just handling/formatting the data correctly. I get it that this is an important aspect (maybe even *the* most important aspect) of data analysis in the real world, but here, it distracts from the real challenge, model building.\nI'd be interested in your thoughts on this.",
      "votes": null
    },
    {
      "id": "1116205",
      "postDate": "12/17/2020 00:27:21",
      "content": "<p>Welcome to the reality of data science.<br>\nAnd to be fair, RLE string is much more space-efficient than npz files or even png masks, and TIFF doesn't compress images too much. There's always a reason</p>\n<p>And that's just about images. If you're interested in speech recognition/audio detection, pre-processing plays an important role. It would be nice if kaggle provides processed MFCCs for human speech, but what if a kaggler wants to use mel-spectrogram, or just the original waveform files to build models?</p>",
      "rawMarkdown": "Welcome to the reality of data science.\nAnd to be fair, RLE string is much more space-efficient than npz files or even png masks, and TIFF doesn't compress images too much. There's always a reason\n\nAnd that's just about images. If you're interested in speech recognition/audio detection, pre-processing plays an important role. It would be nice if kaggle provides processed MFCCs for human speech, but what if a kaggler wants to use mel-spectrogram, or just the original waveform files to build models?",
      "votes": null
    },
    {
      "id": "1116883",
      "postDate": "12/17/2020 14:39:35",
      "content": "<p>Yeah, you're absolutely right about the \"real world\". I guess I thought more from the organizer's point of view. If I were to host a competition, I wouldn't want people to design several (naturally incompatible) data pipelines, I would want them to develop models that I could implement in my own pipeline easily. But maybe they are more interested in pipelines. At least this competition has the clearest explanation of a TFRecords pipeline that I've seen on Kaggle… 😉</p>",
      "rawMarkdown": "Yeah, you're absolutely right about the \"real world\". I guess I thought more from the organizer's point of view. If I were to host a competition, I wouldn't want people to design several (naturally incompatible) data pipelines, I would want them to develop models that I could implement in my own pipeline easily. But maybe they are more interested in pipelines. At least this competition has the clearest explanation of a TFRecords pipeline that I've seen on Kaggle... 😉",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1116205,
      "author_name": "jihangz",
      "author_url": "",
      "post_date": "12/17/2020 00:27:21",
      "content": "<p>Welcome to the reality of data science.<br>\nAnd to be fair, RLE string is much more space-efficient than npz files or even png masks, and TIFF doesn't compress images too much. There's always a reason</p>\n<p>And that's just about images. If you're interested in speech recognition/audio detection, pre-processing plays an important role. It would be nice if kaggle provides processed MFCCs for human speech, but what if a kaggler wants to use mel-spectrogram, or just the original waveform files to build models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1116883,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "12/17/2020 14:39:35",
          "content": "<p>Yeah, you're absolutely right about the \"real world\". I guess I thought more from the organizer's point of view. If I were to host a competition, I wouldn't want people to design several (naturally incompatible) data pipelines, I would want them to develop models that I could implement in my own pipeline easily. But maybe they are more interested in pipelines. At least this competition has the clearest explanation of a TFRecords pipeline that I've seen on Kaggle… 😉</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1115507": "Might be a little off-topic, but after facing the same problem in different competitions over time, I've started wondering why the data formats in kaggle competitions are often so (unnecessarily?) complicated. [These two functions](https://www.kaggle.com/friedchips/fully-correct-hubmap-rle-encoding-and-decoding) cost me a full evening...\nLike in this competition: TIFFs in different formats, submissions as RLE strings. Why? Compressed numpy arrays (npz) would work just fine in both cases.\nA lot of time and effort is spent on just handling/formatting the data correctly. I get it that this is an important aspect (maybe even *the* most important aspect) of data analysis in the real world, but here, it distracts from the real challenge, model building.\nI'd be interested in your thoughts on this.",
    "1116205": "Welcome to the reality of data science.\nAnd to be fair, RLE string is much more space-efficient than npz files or even png masks, and TIFF doesn't compress images too much. There's always a reason\n\nAnd that's just about images. If you're interested in speech recognition/audio detection, pre-processing plays an important role. It would be nice if kaggle provides processed MFCCs for human speech, but what if a kaggler wants to use mel-spectrogram, or just the original waveform files to build models?",
    "1116883": "Yeah, you're absolutely right about the \"real world\". I guess I thought more from the organizer's point of view. If I were to host a competition, I wouldn't want people to design several (naturally incompatible) data pipelines, I would want them to develop models that I could implement in my own pipeline easily. But maybe they are more interested in pipelines. At least this competition has the clearest explanation of a TFRecords pipeline that I've seen on Kaggle... 😉"
  },
  "source": "meta"
}