{
  "id": 200460,
  "title": "A way to scrape more data from google",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/200460",
  "author_name": "",
  "post_date": "2020-11-30T17:03:19.557036700Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I want to share with you another way to scrape more data to feed into your models. Here are the steps to reproduce:</p>\n<h3>Steps</h3>\n<ol>\n<li>Go to google images and type something like <a href=\"https://www.google.com/search?q=Mosaic+Disease&amp;tbm=isch&amp;ved=2ahUKEwiQ56vr3KrtAhWQzSoKHe6OAH4Q2-cCegQIABAA&amp;oq=Mosaic+Disease&amp;gs_lcp=CgNpbWcQAzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQE1D5FFid6gFgmu0BaABwAHgAgAFHiAFHkgEBMZgBAKABAaoBC2d3cy13aXotaW1nsAEAwAEB&amp;sclient=img&amp;ei=viLFX5CQLZCbqwHunYLwBw&amp;bih=937&amp;biw=1920\" target=\"_blank\">Mosaic Disease</a>. Scroll down a little bit. Remember - you will be able to download only those images that you will see on this page. The more you scroll - the more data you will donwload later.</li>\n<li>Open browser console. You can do it by pressing Ctrl + Shift + I in chrome on windows. Then click on \"Console\" tab.</li>\n<li>Enter the following JavaScript code, which would download a text file with URLs to the images:<br>\n<code>urls=Array.from(document.querySelectorAll('.rg_i')).map(el=&gt; el.hasAttribute('data-src')?el.getAttribute('data-src'):el.getAttribute('data-iurl'));\nwindow.open('data:text/csv;charset=utf-8,' + escape(urls.join('\\n')));\n</code></li>\n<li>Save this file with a name like <code>urls.txt</code></li>\n<li>Run the following python code:</li>\n</ol>\n<pre><code>import requests\n\nwith open('urls.txt', 'r') as urlfile:\n    urls = urlfile.read().split()\n\nfor i, url in enumerate(urls):\n    with open(f'image_{i}.jpg', 'wb') as handle:\n        response = requests.get(url, stream=True)\n\n        for block in response.iter_content(1024):\n            if not block:\n                break\n            handle.write(block)\n</code></pre>\n<h3>Pros</h3>\n<ul>\n<li>Now you have additional data.</li>\n</ul>\n<h3>Cons</h3>\n<ul>\n<li>Be aware that this data might be noisy</li>\n<li>Sizes of the images are relatively small (200x300 or so) since we are downloading a previews.</li>\n</ul>",
  "messages": [
    {
      "id": "1096653",
      "postDate": "11/30/2020 17:03:19",
      "content": "<p>I want to share with you another way to scrape more data to feed into your models. Here are the steps to reproduce:</p>\n<h3>Steps</h3>\n<ol>\n<li>Go to google images and type something like <a href=\"https://www.google.com/search?q=Mosaic+Disease&amp;tbm=isch&amp;ved=2ahUKEwiQ56vr3KrtAhWQzSoKHe6OAH4Q2-cCegQIABAA&amp;oq=Mosaic+Disease&amp;gs_lcp=CgNpbWcQAzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQE1D5FFid6gFgmu0BaABwAHgAgAFHiAFHkgEBMZgBAKABAaoBC2d3cy13aXotaW1nsAEAwAEB&amp;sclient=img&amp;ei=viLFX5CQLZCbqwHunYLwBw&amp;bih=937&amp;biw=1920\" target=\"_blank\">Mosaic Disease</a>. Scroll down a little bit. Remember - you will be able to download only those images that you will see on this page. The more you scroll - the more data you will donwload later.</li>\n<li>Open browser console. You can do it by pressing Ctrl + Shift + I in chrome on windows. Then click on \"Console\" tab.</li>\n<li>Enter the following JavaScript code, which would download a text file with URLs to the images:<br>\n<code>urls=Array.from(document.querySelectorAll('.rg_i')).map(el=&gt; el.hasAttribute('data-src')?el.getAttribute('data-src'):el.getAttribute('data-iurl'));\nwindow.open('data:text/csv;charset=utf-8,' + escape(urls.join('\\n')));\n</code></li>\n<li>Save this file with a name like <code>urls.txt</code></li>\n<li>Run the following python code:</li>\n</ol>\n<pre><code>import requests\n\nwith open('urls.txt', 'r') as urlfile:\n    urls = urlfile.read().split()\n\nfor i, url in enumerate(urls):\n    with open(f'image_{i}.jpg', 'wb') as handle:\n        response = requests.get(url, stream=True)\n\n        for block in response.iter_content(1024):\n            if not block:\n                break\n            handle.write(block)\n</code></pre>\n<h3>Pros</h3>\n<ul>\n<li>Now you have additional data.</li>\n</ul>\n<h3>Cons</h3>\n<ul>\n<li>Be aware that this data might be noisy</li>\n<li>Sizes of the images are relatively small (200x300 or so) since we are downloading a previews.</li>\n</ul>",
      "rawMarkdown": "I want to share with you another way to scrape more data to feed into your models. Here are the steps to reproduce:\n\n### Steps\n1. Go to google images and type something like [Mosaic Disease](https://www.google.com/search?q=Mosaic+Disease&tbm=isch&ved=2ahUKEwiQ56vr3KrtAhWQzSoKHe6OAH4Q2-cCegQIABAA&oq=Mosaic+Disease&gs_lcp=CgNpbWcQAzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQE1D5FFid6gFgmu0BaABwAHgAgAFHiAFHkgEBMZgBAKABAaoBC2d3cy13aXotaW1nsAEAwAEB&sclient=img&ei=viLFX5CQLZCbqwHunYLwBw&bih=937&biw=1920). Scroll down a little bit. Remember - you will be able to download only those images that you will see on this page. The more you scroll - the more data you will donwload later.\n2. Open browser console. You can do it by pressing Ctrl + Shift + I in chrome on windows. Then click on \"Console\" tab.\n3. Enter the following JavaScript code, which would download a text file with URLs to the images:\n`urls=Array.from(document.querySelectorAll('.rg_i')).map(el=> el.hasAttribute('data-src')?el.getAttribute('data-src'):el.getAttribute('data-iurl'));\nwindow.open('data:text/csv;charset=utf-8,' + escape(urls.join('\\n')));\n`\n4. Save this file with a name like `urls.txt`\n5. Run the following python code:\n```\nimport requests\n\nwith open('urls.txt', 'r') as urlfile:\n    urls = urlfile.read().split()\n\nfor i, url in enumerate(urls):\n    with open(f'image_{i}.jpg', 'wb') as handle:\n        response = requests.get(url, stream=True)\n\n        for block in response.iter_content(1024):\n            if not block:\n                break\n            handle.write(block)\n```\n\n### Pros\n* Now you have additional data.\n\n### Cons\n* Be aware that this data might be noisy\n* Sizes of the images are relatively small (200x300 or so) since we are downloading a previews.",
      "votes": null
    },
    {
      "id": "1097643",
      "postDate": "12/01/2020 06:48:25",
      "content": "<p>Be careful with this approach, external data needs to be accessible to all the other competitors with no additional cost according to rules. So you may need to share the data you created here on Kaggle, although I am not 100% sure and still waiting for Kaggle team to reply. Expanding the dataset is definitely way to go!</p>",
      "rawMarkdown": "Be careful with this approach, external data needs to be accessible to all the other competitors with no additional cost according to rules. So you may need to share the data you created here on Kaggle, although I am not 100% sure and still waiting for Kaggle team to reply. Expanding the dataset is definitely way to go!",
      "votes": null
    },
    {
      "id": "1098278",
      "postDate": "12/01/2020 14:42:42",
      "content": "<p>I think in a previous competition the issue with scraped data is you can't show you have the right to use the data for this competition. Even images that are on the internet might still officially have limitations in their use.</p>",
      "rawMarkdown": "I think in a previous competition the issue with scraped data is you can't show you have the right to use the data for this competition. Even images that are on the internet might still officially have limitations in their use.",
      "votes": null
    },
    {
      "id": "1098291",
      "postDate": "12/01/2020 14:51:34",
      "content": "<p>You can choose a license of the images in search settings.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1696976%2F77f3e5cd94bd8678468fe0967c0d5a02%2F.png?generation=1606834282678200&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "You can choose a license of the images in search settings.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1696976%2F77f3e5cd94bd8678468fe0967c0d5a02%2F.png?generation=1606834282678200&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1097643,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "12/01/2020 06:48:25",
      "content": "<p>Be careful with this approach, external data needs to be accessible to all the other competitors with no additional cost according to rules. So you may need to share the data you created here on Kaggle, although I am not 100% sure and still waiting for Kaggle team to reply. Expanding the dataset is definitely way to go!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1098278,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "12/01/2020 14:42:42",
      "content": "<p>I think in a previous competition the issue with scraped data is you can't show you have the right to use the data for this competition. Even images that are on the internet might still officially have limitations in their use.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1098291,
          "author_name": "nroman",
          "author_url": "",
          "post_date": "12/01/2020 14:51:34",
          "content": "<p>You can choose a license of the images in search settings.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1696976%2F77f3e5cd94bd8678468fe0967c0d5a02%2F.png?generation=1606834282678200&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1096653": "I want to share with you another way to scrape more data to feed into your models. Here are the steps to reproduce:\n\n### Steps\n1. Go to google images and type something like [Mosaic Disease](https://www.google.com/search?q=Mosaic+Disease&tbm=isch&ved=2ahUKEwiQ56vr3KrtAhWQzSoKHe6OAH4Q2-cCegQIABAA&oq=Mosaic+Disease&gs_lcp=CgNpbWcQAzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQEzIECAAQE1D5FFid6gFgmu0BaABwAHgAgAFHiAFHkgEBMZgBAKABAaoBC2d3cy13aXotaW1nsAEAwAEB&sclient=img&ei=viLFX5CQLZCbqwHunYLwBw&bih=937&biw=1920). Scroll down a little bit. Remember - you will be able to download only those images that you will see on this page. The more you scroll - the more data you will donwload later.\n2. Open browser console. You can do it by pressing Ctrl + Shift + I in chrome on windows. Then click on \"Console\" tab.\n3. Enter the following JavaScript code, which would download a text file with URLs to the images:\n`urls=Array.from(document.querySelectorAll('.rg_i')).map(el=> el.hasAttribute('data-src')?el.getAttribute('data-src'):el.getAttribute('data-iurl'));\nwindow.open('data:text/csv;charset=utf-8,' + escape(urls.join('\\n')));\n`\n4. Save this file with a name like `urls.txt`\n5. Run the following python code:\n```\nimport requests\n\nwith open('urls.txt', 'r') as urlfile:\n    urls = urlfile.read().split()\n\nfor i, url in enumerate(urls):\n    with open(f'image_{i}.jpg', 'wb') as handle:\n        response = requests.get(url, stream=True)\n\n        for block in response.iter_content(1024):\n            if not block:\n                break\n            handle.write(block)\n```\n\n### Pros\n* Now you have additional data.\n\n### Cons\n* Be aware that this data might be noisy\n* Sizes of the images are relatively small (200x300 or so) since we are downloading a previews.",
    "1097643": "Be careful with this approach, external data needs to be accessible to all the other competitors with no additional cost according to rules. So you may need to share the data you created here on Kaggle, although I am not 100% sure and still waiting for Kaggle team to reply. Expanding the dataset is definitely way to go!",
    "1098278": "I think in a previous competition the issue with scraped data is you can't show you have the right to use the data for this competition. Even images that are on the internet might still officially have limitations in their use.",
    "1098291": "You can choose a license of the images in search settings.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1696976%2F77f3e5cd94bd8678468fe0967c0d5a02%2F.png?generation=1606834282678200&alt=media)"
  },
  "source": "meta"
}