{
  "id": 223230,
  "title": "Low quality of images",
  "url": "/competitions/bms-molecular-translation/discussion/223230",
  "author_name": "",
  "post_date": "2021-03-03T02:42:05.156627100Z",
  "votes": 9,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Documents scanned contemporarily would produce much higher resolution, grayscale (or color) images as opposed to low resolution, bi-level, fax quality images provided. My solution, as well as many real world solutions, would greatly benefit from higher resolution, grayscale images. Is there any reason for this image format selection, other than tailoring the data to a typical DNN, which struggles with high resolution images?</p>",
  "messages": [
    {
      "id": "1224733",
      "postDate": "03/03/2021 02:42:05",
      "content": "<p>Documents scanned contemporarily would produce much higher resolution, grayscale (or color) images as opposed to low resolution, bi-level, fax quality images provided. My solution, as well as many real world solutions, would greatly benefit from higher resolution, grayscale images. Is there any reason for this image format selection, other than tailoring the data to a typical DNN, which struggles with high resolution images?</p>",
      "rawMarkdown": "Documents scanned contemporarily would produce much higher resolution, grayscale (or color) images as opposed to low resolution, bi-level, fax quality images provided. My solution, as well as many real world solutions, would greatly benefit from higher resolution, grayscale images. Is there any reason for this image format selection, other than tailoring the data to a typical DNN, which struggles with high resolution images?",
      "votes": null
    },
    {
      "id": "1225080",
      "postDate": "03/03/2021 09:40:15",
      "content": "<p>I also noticed that when browsing through the images, they are of really, really low quality.</p>",
      "rawMarkdown": "I also noticed that when browsing through the images, they are of really, really low quality.",
      "votes": null
    },
    {
      "id": "1225364",
      "postDate": "03/03/2021 15:07:04",
      "content": "<p>I suppose this is to be considered a part of the problem to be taken care of?</p>",
      "rawMarkdown": "I suppose this is to be considered a part of the problem to be taken care of?",
      "votes": null
    },
    {
      "id": "1225635",
      "postDate": "03/03/2021 19:05:07",
      "content": "<p>It would be, if the fax machine was an exclusive method of communicating structural formulas. If you look at the big picture, the goal is to have a piece of software performing well in the real world, in order to \"help chemists expand access to collective chemical research\". A somewhat unrealistic set of synthetic images may not advance this goal well. Deep learning is full of models performing well on synthetic or otherwise simplified training sets and failing miserably in the real world.</p>\n<p>I propose that organizers publish a companion data set of higher resolution grayscale version of all images. Competitors would be free to choose one or both sets for training and evaluation.</p>",
      "rawMarkdown": "It would be, if the fax machine was an exclusive method of communicating structural formulas. If you look at the big picture, the goal is to have a piece of software performing well in the real world, in order to \"help chemists expand access to collective chemical research\". A somewhat unrealistic set of synthetic images may not advance this goal well. Deep learning is full of models performing well on synthetic or otherwise simplified training sets and failing miserably in the real world.\n\nI propose that organizers publish a companion data set of higher resolution grayscale version of all images. Competitors would be free to choose one or both sets for training and evaluation.",
      "votes": null
    },
    {
      "id": "1225812",
      "postDate": "03/04/2021 01:05:15",
      "content": "<p>From the overview description:  \"Recent publications are also annotated with machine-readable chemical descriptions (InChI), but there are decades of scanned documents that can't be automatically searched for specific chemical depictions. Automated recognition of optical chemical structures, with the help of machine learning, could speed up research and development efforts\"</p>\n<p>Sounds to me like there is a potential wealth of information hidden in older images that were taken or produced before InChI was even a thing.  Chemists would have to painstakingly go through those old images and label them by hand but that's exactly the type of boring, repetitive task that a smart computer program should handle for them</p>",
      "rawMarkdown": "From the overview description:  \"Recent publications are also annotated with machine-readable chemical descriptions (InChI), but there are decades of scanned documents that can't be automatically searched for specific chemical depictions. Automated recognition of optical chemical structures, with the help of machine learning, could speed up research and development efforts\"\n\nSounds to me like there is a potential wealth of information hidden in older images that were taken or produced before InChI was even a thing.  Chemists would have to painstakingly go through those old images and label them by hand but that's exactly the type of boring, repetitive task that a smart computer program should handle for them",
      "votes": null
    },
    {
      "id": "1225826",
      "postDate": "03/04/2021 01:39:15",
      "content": "<p>This makes sense to me, definitely sounds like a goal worth pursuing :)</p>",
      "rawMarkdown": "This makes sense to me, definitely sounds like a goal worth pursuing :)",
      "votes": null
    },
    {
      "id": "1225836",
      "postDate": "03/04/2021 01:55:23",
      "content": "<p>No disagreement here. What I'm arguing is that higher resolution grayscale scanners were commonly available for almost four decades now. Taking into account the high error rate of current OCSR, any improvement gained by better quality input would save a lot of unnecessary manual transcription labor. There may be some cases where the only available image is of similar quality to the data offered here, but I suspect it is a minority of cases. Reducing the quality of a whole dataset to the lowest common denominator is counterproductive for real world application.</p>",
      "rawMarkdown": "No disagreement here. What I'm arguing is that higher resolution grayscale scanners were commonly available for almost four decades now. Taking into account the high error rate of current OCSR, any improvement gained by better quality input would save a lot of unnecessary manual transcription labor. There may be some cases where the only available image is of similar quality to the data offered here, but I suspect it is a minority of cases. Reducing the quality of a whole dataset to the lowest common denominator is counterproductive for real world application.",
      "votes": null
    },
    {
      "id": "1225841",
      "postDate": "03/04/2021 02:02:38",
      "content": "<p><a href=\"https://www.kaggle.com/pauljurczak\" target=\"_blank\">@pauljurczak</a> from what I understand, according to you there is no particular reason for the quality of images which is also very much possible</p>",
      "rawMarkdown": "pauljurczak from what I understand, according to you there is no particular reason for the quality of images which is also very much possible",
      "votes": null
    },
    {
      "id": "1226496",
      "postDate": "03/04/2021 15:28:25",
      "content": "<p>Also the resolution varies greatly in the train and the test data set both in x- and y-direction… </p>",
      "rawMarkdown": "Also the resolution varies greatly in the train and the test data set both in x- and y-direction...",
      "votes": null
    },
    {
      "id": "1226927",
      "postDate": "03/05/2021 03:07:02",
      "content": "<p>\"There may be some cases where the only available image is of similar quality to the data offered here, but I suspect it is a minority of cases\" - I don't know if I can agree with this assumption.  From what I understand (and mind you, I'm no real chemist and have probably learned more about molecules in the past couple days than probably my entire life, so take this with a grain of salt), but from what I understand there are infinite atomic configurations you can string together to form molecules.  If that is the case and this dataset has ~4 million images, then we could reasonably suspect that those 4 million images might only represent a tiny fraction of all of the molecular structures ever drawn up by chemists all around the world</p>\n<p>Better equipment for imaging might have been around for some time, but did all chem labs have access to that and did they even bother to upload their files in the highest resolution?  They might have never suspected that one day a bunch of Kagglers would come together and try to run ML algorithms on their pictures and they might've just assumed that the quality they uploaded and saved would be good enough because a human, especially one trained in chemistry, could still read and interpret these pictures.  Even if you had some forward thinking chemists that made sure to save all of their files properly, as the competition overview mentions: \"Historical sources often have some level of image corruption, which reduces performance to near zero\".  I would not be surprised if a pharmaceutical company like BMS has a mountain of old pictures of molecules, some of which may have never been indexed anywhere in a connected database.  Heck, how would you even know that those images aren't already saved somewhere in better quality if those old images don't have their corresponding InChI's saved alongside them?  </p>\n<p>Also (and sorry for the long post, this is just genuinely fun to think about), but have you thought about using ML/AI to increase the quality of the pictures?  If Peter Jackson can turn black and white WWI footage into a color movie then maybe we can also devise ways to make these images better.  Okay - that's probably oversimplified, but seriously, maybe just take a couple of crisp images and run them through lossy compression algorithms.  That could give you your features and your labels for training</p>",
      "rawMarkdown": "\"There may be some cases where the only available image is of similar quality to the data offered here, but I suspect it is a minority of cases\" - I don't know if I can agree with this assumption.  From what I understand (and mind you, I'm no real chemist and have probably learned more about molecules in the past couple days than probably my entire life, so take this with a grain of salt), but from what I understand there are infinite atomic configurations you can string together to form molecules.  If that is the case and this dataset has ~4 million images, then we could reasonably suspect that those 4 million images might only represent a tiny fraction of all of the molecular structures ever drawn up by chemists all around the world\n\nBetter equipment for imaging might have been around for some time, but did all chem labs have access to that and did they even bother to upload their files in the highest resolution?  They might have never suspected that one day a bunch of Kagglers would come together and try to run ML algorithms on their pictures and they might've just assumed that the quality they uploaded and saved would be good enough because a human, especially one trained in chemistry, could still read and interpret these pictures.  Even if you had some forward thinking chemists that made sure to save all of their files properly, as the competition overview mentions: \"Historical sources often have some level of image corruption, which reduces performance to near zero\".  I would not be surprised if a pharmaceutical company like BMS has a mountain of old pictures of molecules, some of which may have never been indexed anywhere in a connected database.  Heck, how would you even know that those images aren't already saved somewhere in better quality if those old images don't have their corresponding InChI's saved alongside them?  \n\nAlso (and sorry for the long post, this is just genuinely fun to think about), but have you thought about using ML/AI to increase the quality of the pictures?  If Peter Jackson can turn black and white WWI footage into a color movie then maybe we can also devise ways to make these images better.  Okay - that's probably oversimplified, but seriously, maybe just take a couple of crisp images and run them through lossy compression algorithms.  That could give you your features and your labels for training",
      "votes": null
    },
    {
      "id": "1226967",
      "postDate": "03/05/2021 04:17:32",
      "content": "<blockquote>\n  <p>did all chem labs have access to that and did they even bother to upload their files in the highest resolution</p>\n</blockquote>\n<p>I'm assuming that most of the real world images come from scanned publications, which usually assures pretty good quality. Granted, some internal documents may be of low quality or unreadable. This is not how the data for this competion was created, though.</p>\n<p>My post is about calibrating the quality level of synthetic images to match the real world better. From the overview (<a href=\"https://www.kaggle.com/c/bms-molecular-translation/overview):\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/overview):</a> \"large set of synthetic image data generated by Bristol-Myers Squibb\". Organizers of this competion can dial the resolution, noise level, pixel depth and other parameters of these images. They chose one of great number of options available to them.</p>\n<p>The point is moot anyway, because I don't think Kaggle ever changed the data in a significant way after the competion started.  </p>",
      "rawMarkdown": "> did all chem labs have access to that and did they even bother to upload their files in the highest resolution\n\nI'm assuming that most of the real world images come from scanned publications, which usually assures pretty good quality. Granted, some internal documents may be of low quality or unreadable. This is not how the data for this competion was created, though.\n\nMy post is about calibrating the quality level of synthetic images to match the real world better. From the overview (https://www.kaggle.com/c/bms-molecular-translation/overview): \"large set of synthetic image data generated by Bristol-Myers Squibb\". Organizers of this competion can dial the resolution, noise level, pixel depth and other parameters of these images. They chose one of great number of options available to them.\n\nThe point is moot anyway, because I don't think Kaggle ever changed the data in a significant way after the competion started.",
      "votes": null
    },
    {
      "id": "1230689",
      "postDate": "03/08/2021 11:09:20",
      "content": "<p><a href=\"https://www.kaggle.com/leonmcg\" target=\"_blank\">@leonmcg</a> &gt;  maybe we can also devise ways to make these images better. Okay - that's probably oversimplified</p>\n<p>I'm going to experiment with SuperRes, not sure if that would work nicely given the \"graininess\" of the images but might be worth an experiment. </p>\n<p>I suspect doing the opposite as you suggested might be a better effort -&gt; Using images from other comps of higher res and adding noise to these.</p>",
      "rawMarkdown": "leonmcg >  maybe we can also devise ways to make these images better. Okay - that's probably oversimplified\n\nI'm going to experiment with SuperRes, not sure if that would work nicely given the \"graininess\" of the images but might be worth an experiment. \n\nI suspect doing the opposite as you suggested might be a better effort -> Using images from other comps of higher res and adding noise to these.",
      "votes": null
    },
    {
      "id": "1231323",
      "postDate": "03/08/2021 21:54:58",
      "content": "<p>Yes, the images are quite bad, but they are engineered to reflect a scenario where documents have been microfilmed then digitized.   In my experience, fax quality is most common in this case.  Sadly, time and cost have outweighed fidelity for many records.</p>",
      "rawMarkdown": "Yes, the images are quite bad, but they are engineered to reflect a scenario where documents have been microfilmed then digitized.   In my experience, fax quality is most common in this case.  Sadly, time and cost have outweighed fidelity for many records.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1225080,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "03/03/2021 09:40:15",
      "content": "<p>I also noticed that when browsing through the images, they are of really, really low quality.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1225364,
      "author_name": "arka47",
      "author_url": "",
      "post_date": "03/03/2021 15:07:04",
      "content": "<p>I suppose this is to be considered a part of the problem to be taken care of?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1225635,
          "author_name": "pauljurczak",
          "author_url": "",
          "post_date": "03/03/2021 19:05:07",
          "content": "<p>It would be, if the fax machine was an exclusive method of communicating structural formulas. If you look at the big picture, the goal is to have a piece of software performing well in the real world, in order to \"help chemists expand access to collective chemical research\". A somewhat unrealistic set of synthetic images may not advance this goal well. Deep learning is full of models performing well on synthetic or otherwise simplified training sets and failing miserably in the real world.</p>\n<p>I propose that organizers publish a companion data set of higher resolution grayscale version of all images. Competitors would be free to choose one or both sets for training and evaluation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225812,
          "author_name": "leonmcg",
          "author_url": "",
          "post_date": "03/04/2021 01:05:15",
          "content": "<p>From the overview description:  \"Recent publications are also annotated with machine-readable chemical descriptions (InChI), but there are decades of scanned documents that can't be automatically searched for specific chemical depictions. Automated recognition of optical chemical structures, with the help of machine learning, could speed up research and development efforts\"</p>\n<p>Sounds to me like there is a potential wealth of information hidden in older images that were taken or produced before InChI was even a thing.  Chemists would have to painstakingly go through those old images and label them by hand but that's exactly the type of boring, repetitive task that a smart computer program should handle for them</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225826,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/04/2021 01:39:15",
          "content": "<p>This makes sense to me, definitely sounds like a goal worth pursuing :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225836,
          "author_name": "pauljurczak",
          "author_url": "",
          "post_date": "03/04/2021 01:55:23",
          "content": "<p>No disagreement here. What I'm arguing is that higher resolution grayscale scanners were commonly available for almost four decades now. Taking into account the high error rate of current OCSR, any improvement gained by better quality input would save a lot of unnecessary manual transcription labor. There may be some cases where the only available image is of similar quality to the data offered here, but I suspect it is a minority of cases. Reducing the quality of a whole dataset to the lowest common denominator is counterproductive for real world application.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225841,
          "author_name": "arka47",
          "author_url": "",
          "post_date": "03/04/2021 02:02:38",
          "content": "<p><a href=\"https://www.kaggle.com/pauljurczak\" target=\"_blank\">@pauljurczak</a> from what I understand, according to you there is no particular reason for the quality of images which is also very much possible</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1226927,
          "author_name": "leonmcg",
          "author_url": "",
          "post_date": "03/05/2021 03:07:02",
          "content": "<p>\"There may be some cases where the only available image is of similar quality to the data offered here, but I suspect it is a minority of cases\" - I don't know if I can agree with this assumption.  From what I understand (and mind you, I'm no real chemist and have probably learned more about molecules in the past couple days than probably my entire life, so take this with a grain of salt), but from what I understand there are infinite atomic configurations you can string together to form molecules.  If that is the case and this dataset has ~4 million images, then we could reasonably suspect that those 4 million images might only represent a tiny fraction of all of the molecular structures ever drawn up by chemists all around the world</p>\n<p>Better equipment for imaging might have been around for some time, but did all chem labs have access to that and did they even bother to upload their files in the highest resolution?  They might have never suspected that one day a bunch of Kagglers would come together and try to run ML algorithms on their pictures and they might've just assumed that the quality they uploaded and saved would be good enough because a human, especially one trained in chemistry, could still read and interpret these pictures.  Even if you had some forward thinking chemists that made sure to save all of their files properly, as the competition overview mentions: \"Historical sources often have some level of image corruption, which reduces performance to near zero\".  I would not be surprised if a pharmaceutical company like BMS has a mountain of old pictures of molecules, some of which may have never been indexed anywhere in a connected database.  Heck, how would you even know that those images aren't already saved somewhere in better quality if those old images don't have their corresponding InChI's saved alongside them?  </p>\n<p>Also (and sorry for the long post, this is just genuinely fun to think about), but have you thought about using ML/AI to increase the quality of the pictures?  If Peter Jackson can turn black and white WWI footage into a color movie then maybe we can also devise ways to make these images better.  Okay - that's probably oversimplified, but seriously, maybe just take a couple of crisp images and run them through lossy compression algorithms.  That could give you your features and your labels for training</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1226967,
          "author_name": "pauljurczak",
          "author_url": "",
          "post_date": "03/05/2021 04:17:32",
          "content": "<blockquote>\n  <p>did all chem labs have access to that and did they even bother to upload their files in the highest resolution</p>\n</blockquote>\n<p>I'm assuming that most of the real world images come from scanned publications, which usually assures pretty good quality. Granted, some internal documents may be of low quality or unreadable. This is not how the data for this competion was created, though.</p>\n<p>My post is about calibrating the quality level of synthetic images to match the real world better. From the overview (<a href=\"https://www.kaggle.com/c/bms-molecular-translation/overview):\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/overview):</a> \"large set of synthetic image data generated by Bristol-Myers Squibb\". Organizers of this competion can dial the resolution, noise level, pixel depth and other parameters of these images. They chose one of great number of options available to them.</p>\n<p>The point is moot anyway, because I don't think Kaggle ever changed the data in a significant way after the competion started.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1230689,
          "author_name": "init27",
          "author_url": "",
          "post_date": "03/08/2021 11:09:20",
          "content": "<p><a href=\"https://www.kaggle.com/leonmcg\" target=\"_blank\">@leonmcg</a> &gt;  maybe we can also devise ways to make these images better. Okay - that's probably oversimplified</p>\n<p>I'm going to experiment with SuperRes, not sure if that would work nicely given the \"graininess\" of the images but might be worth an experiment. </p>\n<p>I suspect doing the opposite as you suggested might be a better effort -&gt; Using images from other comps of higher res and adding noise to these.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1226496,
      "author_name": "erichahn",
      "author_url": "",
      "post_date": "03/04/2021 15:28:25",
      "content": "<p>Also the resolution varies greatly in the train and the test data set both in x- and y-direction… </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1231323,
      "author_name": "jakealbrecht1337",
      "author_url": "",
      "post_date": "03/08/2021 21:54:58",
      "content": "<p>Yes, the images are quite bad, but they are engineered to reflect a scenario where documents have been microfilmed then digitized.   In my experience, fax quality is most common in this case.  Sadly, time and cost have outweighed fidelity for many records.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1224733": "Documents scanned contemporarily would produce much higher resolution, grayscale (or color) images as opposed to low resolution, bi-level, fax quality images provided. My solution, as well as many real world solutions, would greatly benefit from higher resolution, grayscale images. Is there any reason for this image format selection, other than tailoring the data to a typical DNN, which struggles with high resolution images?",
    "1225080": "I also noticed that when browsing through the images, they are of really, really low quality.",
    "1225364": "I suppose this is to be considered a part of the problem to be taken care of?",
    "1225635": "It would be, if the fax machine was an exclusive method of communicating structural formulas. If you look at the big picture, the goal is to have a piece of software performing well in the real world, in order to \"help chemists expand access to collective chemical research\". A somewhat unrealistic set of synthetic images may not advance this goal well. Deep learning is full of models performing well on synthetic or otherwise simplified training sets and failing miserably in the real world.\n\nI propose that organizers publish a companion data set of higher resolution grayscale version of all images. Competitors would be free to choose one or both sets for training and evaluation.",
    "1225812": "From the overview description:  \"Recent publications are also annotated with machine-readable chemical descriptions (InChI), but there are decades of scanned documents that can't be automatically searched for specific chemical depictions. Automated recognition of optical chemical structures, with the help of machine learning, could speed up research and development efforts\"\n\nSounds to me like there is a potential wealth of information hidden in older images that were taken or produced before InChI was even a thing.  Chemists would have to painstakingly go through those old images and label them by hand but that's exactly the type of boring, repetitive task that a smart computer program should handle for them",
    "1225826": "This makes sense to me, definitely sounds like a goal worth pursuing :)",
    "1225836": "No disagreement here. What I'm arguing is that higher resolution grayscale scanners were commonly available for almost four decades now. Taking into account the high error rate of current OCSR, any improvement gained by better quality input would save a lot of unnecessary manual transcription labor. There may be some cases where the only available image is of similar quality to the data offered here, but I suspect it is a minority of cases. Reducing the quality of a whole dataset to the lowest common denominator is counterproductive for real world application.",
    "1225841": "pauljurczak from what I understand, according to you there is no particular reason for the quality of images which is also very much possible",
    "1226496": "Also the resolution varies greatly in the train and the test data set both in x- and y-direction...",
    "1226927": "\"There may be some cases where the only available image is of similar quality to the data offered here, but I suspect it is a minority of cases\" - I don't know if I can agree with this assumption.  From what I understand (and mind you, I'm no real chemist and have probably learned more about molecules in the past couple days than probably my entire life, so take this with a grain of salt), but from what I understand there are infinite atomic configurations you can string together to form molecules.  If that is the case and this dataset has ~4 million images, then we could reasonably suspect that those 4 million images might only represent a tiny fraction of all of the molecular structures ever drawn up by chemists all around the world\n\nBetter equipment for imaging might have been around for some time, but did all chem labs have access to that and did they even bother to upload their files in the highest resolution?  They might have never suspected that one day a bunch of Kagglers would come together and try to run ML algorithms on their pictures and they might've just assumed that the quality they uploaded and saved would be good enough because a human, especially one trained in chemistry, could still read and interpret these pictures.  Even if you had some forward thinking chemists that made sure to save all of their files properly, as the competition overview mentions: \"Historical sources often have some level of image corruption, which reduces performance to near zero\".  I would not be surprised if a pharmaceutical company like BMS has a mountain of old pictures of molecules, some of which may have never been indexed anywhere in a connected database.  Heck, how would you even know that those images aren't already saved somewhere in better quality if those old images don't have their corresponding InChI's saved alongside them?  \n\nAlso (and sorry for the long post, this is just genuinely fun to think about), but have you thought about using ML/AI to increase the quality of the pictures?  If Peter Jackson can turn black and white WWI footage into a color movie then maybe we can also devise ways to make these images better.  Okay - that's probably oversimplified, but seriously, maybe just take a couple of crisp images and run them through lossy compression algorithms.  That could give you your features and your labels for training",
    "1226967": "> did all chem labs have access to that and did they even bother to upload their files in the highest resolution\n\nI'm assuming that most of the real world images come from scanned publications, which usually assures pretty good quality. Granted, some internal documents may be of low quality or unreadable. This is not how the data for this competion was created, though.\n\nMy post is about calibrating the quality level of synthetic images to match the real world better. From the overview (https://www.kaggle.com/c/bms-molecular-translation/overview): \"large set of synthetic image data generated by Bristol-Myers Squibb\". Organizers of this competion can dial the resolution, noise level, pixel depth and other parameters of these images. They chose one of great number of options available to them.\n\nThe point is moot anyway, because I don't think Kaggle ever changed the data in a significant way after the competion started.",
    "1230689": "leonmcg >  maybe we can also devise ways to make these images better. Okay - that's probably oversimplified\n\nI'm going to experiment with SuperRes, not sure if that would work nicely given the \"graininess\" of the images but might be worth an experiment. \n\nI suspect doing the opposite as you suggested might be a better effort -> Using images from other comps of higher res and adding noise to these.",
    "1231323": "Yes, the images are quite bad, but they are engineered to reflect a scenario where documents have been microfilmed then digitized.   In my experience, fax quality is most common in this case.  Sadly, time and cost have outweighed fidelity for many records."
  },
  "source": "meta"
}