{
  "id": 66575,
  "title": "How do I get more than 1000 images from competition dataset",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/66575",
  "author_name": "",
  "post_date": "2018-09-23T01:38:09.591101100Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, </p>\n\n<p>I am new to Kaggle.  I tried to run a few kernels (e.g. <a href=\"https://www.kaggle.com/kmader/inceptionv3-for-retinopathy-gpu-hr\">https://www.kaggle.com/kmader/inceptionv3-for-retinopathy-gpu-hr</a>) with this competition.</p>\n\n<p>It runs great, but it can only pull 1000 images out of 35000 competition dataset images, how do I configure it to get all the 35000 images in my notebook?</p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "392100",
      "postDate": "09/23/2018 01:38:09",
      "content": "<p>Hi, </p>\n\n<p>I am new to Kaggle.  I tried to run a few kernels (e.g. <a href=\"https://www.kaggle.com/kmader/inceptionv3-for-retinopathy-gpu-hr\">https://www.kaggle.com/kmader/inceptionv3-for-retinopathy-gpu-hr</a>) with this competition.</p>\n\n<p>It runs great, but it can only pull 1000 images out of 35000 competition dataset images, how do I configure it to get all the 35000 images in my notebook?</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Hi, \n\nI am new to Kaggle.  I tried to run a few kernels (e.g. https://www.kaggle.com/kmader/inceptionv3-for-retinopathy-gpu-hr) with this competition.\n\nIt runs great, but it can only pull 1000 images out of 35000 competition dataset images, how do I configure it to get all the 35000 images in my notebook?\n\nThanks.",
      "votes": null
    },
    {
      "id": "398214",
      "postDate": "10/03/2018 17:13:14",
      "content": "<p>endorsed</p>",
      "rawMarkdown": "endorsed",
      "votes": null
    },
    {
      "id": "398215",
      "postDate": "10/03/2018 17:14:20",
      "content": "<p>If you have got your answer please share with me </p>",
      "rawMarkdown": "If you have got your answer please share with me",
      "votes": null
    },
    {
      "id": "463767",
      "postDate": "01/30/2019 15:25:34",
      "content": "<p>You can easily access the 35000 images by simply downloading all the zip file and extracting only train.zip file will gives you all the 35000 images by combining other zip file(001,002,etc) files automatically. After extracting you have all the images and you can access them with only setting the path.</p>\n\n<h1>Here I am building my dataframe (taking patient id  from image coloum and seperating side of image</h1>\n\n<h1>left or right specefying path of image and at last converting my labels into categorical labels)</h1>\n\n<p><strong>temp_df=pd.read_csv('F:\\FYP DATASET\\images\\trainLabels.csv')</strong> # uploading csv to my pandas dataframe\nprint(temp_df.head()) # displaying first 5 objects in dataframe\nimage=temp_df['image'].str.split('_',n=1,expand=True) #splitting Side and Patient ID \ndf = pd.DataFrame()# creating new dataframe object\ndf['eye_side']=image[1] #taking side of Image\ndf['patient_id']=image[0]#taking patient id of an Image</p>\n\n<p><strong>df['path']='F:\\FYP DATASET\\images\\train\\'</strong>#Giving paths of the images \ndf['path']=df['path'].str.cat(temp_df['image']+'.jpeg')#adding Image path and format \ndf['exists'] = df['path'].map(os.path.exists)\ndf=df[df['exists']]\ndf['level']=temp_df['level']# taking levels of Image\ndf['level_cat'] = df['level'].map(lambda x: to_categorical(x, 1+df['level'].max()))#converting my </p>\n\n<h1>labels to categorical_labels</h1>\n\n<p>df</p>\n\n<p>by this code you can access images and exist coloum represents that image is in the specific path. you only need to change the path. You are not able to access this images inside the kaggle because the files are too large you can access in the local machine by using Jupiter notebook</p>",
      "rawMarkdown": "You can easily access the 35000 images by simply downloading all the zip file and extracting only train.zip file will gives you all the 35000 images by combining other zip file(001,002,etc) files automatically. After extracting you have all the images and you can access them with only setting the path.\n\n# Here I am building my dataframe (taking patient id  from image coloum and seperating side of image \n# left or right specefying path of image and at last converting my labels into categorical labels)\n\n\n**temp_df=pd.read_csv('F:\\\\FYP DATASET\\\\images\\\\trainLabels.csv')** # uploading csv to my pandas dataframe\nprint(temp_df.head()) # displaying first 5 objects in dataframe\nimage=temp_df['image'].str.split('_',n=1,expand=True) #splitting Side and Patient ID \ndf = pd.DataFrame()# creating new dataframe object\ndf['eye_side']=image[1] #taking side of Image\ndf['patient_id']=image[0]#taking patient id of an Image\n\n**df['path']='F:\\\\FYP DATASET\\\\images\\\\train\\\\'**#Giving paths of the images \ndf['path']=df['path'].str.cat(temp_df['image']+'.jpeg')#adding Image path and format \ndf['exists'] = df['path'].map(os.path.exists)\ndf=df[df['exists']]\ndf['level']=temp_df['level']# taking levels of Image\ndf['level_cat'] = df['level'].map(lambda x: to_categorical(x, 1+df['level'].max()))#converting my \n\n\n# labels to categorical_labels\ndf\n\nby this code you can access images and exist coloum represents that image is in the specific path. you only need to change the path. You are not able to access this images inside the kaggle because the files are too large you can access in the local machine by using Jupiter notebook",
      "votes": null
    },
    {
      "id": "679120",
      "postDate": "11/22/2019 09:58:06",
      "content": "<p>how to download 35000 images ? I cannot download the file</p>",
      "rawMarkdown": "how to download 35000 images ? I cannot download the file",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 398214,
      "author_name": "aamirk5",
      "author_url": "",
      "post_date": "10/03/2018 17:13:14",
      "content": "<p>endorsed</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 398215,
      "author_name": "aamirk5",
      "author_url": "",
      "post_date": "10/03/2018 17:14:20",
      "content": "<p>If you have got your answer please share with me </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 463767,
      "author_name": "sohaibanwaar1203",
      "author_url": "",
      "post_date": "01/30/2019 15:25:34",
      "content": "<p>You can easily access the 35000 images by simply downloading all the zip file and extracting only train.zip file will gives you all the 35000 images by combining other zip file(001,002,etc) files automatically. After extracting you have all the images and you can access them with only setting the path.</p>\n\n<h1>Here I am building my dataframe (taking patient id  from image coloum and seperating side of image</h1>\n\n<h1>left or right specefying path of image and at last converting my labels into categorical labels)</h1>\n\n<p><strong>temp_df=pd.read_csv('F:\\FYP DATASET\\images\\trainLabels.csv')</strong> # uploading csv to my pandas dataframe\nprint(temp_df.head()) # displaying first 5 objects in dataframe\nimage=temp_df['image'].str.split('_',n=1,expand=True) #splitting Side and Patient ID \ndf = pd.DataFrame()# creating new dataframe object\ndf['eye_side']=image[1] #taking side of Image\ndf['patient_id']=image[0]#taking patient id of an Image</p>\n\n<p><strong>df['path']='F:\\FYP DATASET\\images\\train\\'</strong>#Giving paths of the images \ndf['path']=df['path'].str.cat(temp_df['image']+'.jpeg')#adding Image path and format \ndf['exists'] = df['path'].map(os.path.exists)\ndf=df[df['exists']]\ndf['level']=temp_df['level']# taking levels of Image\ndf['level_cat'] = df['level'].map(lambda x: to_categorical(x, 1+df['level'].max()))#converting my </p>\n\n<h1>labels to categorical_labels</h1>\n\n<p>df</p>\n\n<p>by this code you can access images and exist coloum represents that image is in the specific path. you only need to change the path. You are not able to access this images inside the kaggle because the files are too large you can access in the local machine by using Jupiter notebook</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 679120,
      "author_name": "diviyaprabha14",
      "author_url": "",
      "post_date": "11/22/2019 09:58:06",
      "content": "<p>how to download 35000 images ? I cannot download the file</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "392100": "Hi, \n\nI am new to Kaggle.  I tried to run a few kernels (e.g. https://www.kaggle.com/kmader/inceptionv3-for-retinopathy-gpu-hr) with this competition.\n\nIt runs great, but it can only pull 1000 images out of 35000 competition dataset images, how do I configure it to get all the 35000 images in my notebook?\n\nThanks.",
    "398214": "endorsed",
    "398215": "If you have got your answer please share with me",
    "463767": "You can easily access the 35000 images by simply downloading all the zip file and extracting only train.zip file will gives you all the 35000 images by combining other zip file(001,002,etc) files automatically. After extracting you have all the images and you can access them with only setting the path.\n\n# Here I am building my dataframe (taking patient id  from image coloum and seperating side of image \n# left or right specefying path of image and at last converting my labels into categorical labels)\n\n\n**temp_df=pd.read_csv('F:\\\\FYP DATASET\\\\images\\\\trainLabels.csv')** # uploading csv to my pandas dataframe\nprint(temp_df.head()) # displaying first 5 objects in dataframe\nimage=temp_df['image'].str.split('_',n=1,expand=True) #splitting Side and Patient ID \ndf = pd.DataFrame()# creating new dataframe object\ndf['eye_side']=image[1] #taking side of Image\ndf['patient_id']=image[0]#taking patient id of an Image\n\n**df['path']='F:\\\\FYP DATASET\\\\images\\\\train\\\\'**#Giving paths of the images \ndf['path']=df['path'].str.cat(temp_df['image']+'.jpeg')#adding Image path and format \ndf['exists'] = df['path'].map(os.path.exists)\ndf=df[df['exists']]\ndf['level']=temp_df['level']# taking levels of Image\ndf['level_cat'] = df['level'].map(lambda x: to_categorical(x, 1+df['level'].max()))#converting my \n\n\n# labels to categorical_labels\ndf\n\nby this code you can access images and exist coloum represents that image is in the specific path. you only need to change the path. You are not able to access this images inside the kaggle because the files are too large you can access in the local machine by using Jupiter notebook",
    "679120": "how to download 35000 images ? I cannot download the file"
  },
  "source": "meta"
}