{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":83735,"databundleVersionId":9881586,"sourceType":"competition"},{"sourceId":23812,"sourceType":"datasetVersion","datasetId":17810},{"sourceId":10054098,"sourceType":"datasetVersion","datasetId":6195117},{"sourceId":10054114,"sourceType":"datasetVersion","datasetId":6195128}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Radiology discussions with Gemini-1.5: A long-context AI breakthough\n\nInspired by NotebookLLM. \n\nInitially, goal was to build an x-ray diagnosis LLM using few shot. But then after many trials of few shot learning and generating descriptions on many references, I could not go beyond 50% threshold on the validation data. But then I thought, there are already great techniques, state of the art techniques that can already do that, so why not just try and generate radiology reports based on the xray and final predictions. But then, a radiology report contains much more than images. At this point, I thought well, LLMs might just be good for generating poems. Then during the 5 day course on Gemini, I listened to the podcasts generated by NotebookLLM, and I was amazed at how one can just simulate characters and let them discuss papers in the most human way people. Then, I thought why not build a Multi Model NotebookLLM (research papers + x-rays). I thought shovel all the papers and images into the context window, and then instantly get a response, but nope. Processing documents would take much longer, and this is when i thought about summarizing each of the pages. The idea of summarization came when I tried to upload textbooks (Calculus, Organic Chemistry, Electromagnetism) to Gemini to create a practical guide for self learning. But I kept getting timed out.  I was going to use context caching, but it is a paid feature and I am currently leaving in China, making it impossible to use this. But this below, is my attempt at using Gemini for Long Context application by having two radiologists discuss about medical imaging, pneumonia, medical usage of AI while taking into consideration proper citations of research papers. I have cross checked the final outputs and its stunning. \n\n## Long context strategy for complex problems\n\nI realized that the key wasn't cramming data into the model but feeding it in *structured, summarized chunks*. Here is how it works\n\n- **Summarization**: Instead of uploading entire research papers, the model is prompted to generate concise summaries. Each page becomes a distilled version of its core insights, reducing input size without losing critical information. The steps for summarization are as below\n\n    - **1. Identify the paper title**: First step is to state the title of the paper to establish context\n \n    - **2. Extract title and subtitles**: All titles, subtitles, and headings are extracted and organized hierarchically to reflect the document's structure.\n \n    - **3. Process each title or subtitle**: Provide a detailed summary covering the main ideas, supporting evidence, conclusions, and broader implications. Describe any figures, or tables summarizing key insights and their relevance\n \n    - **4. Self check mechanism** : After processing each section, a self check is performed to ensure that no title or subtitle is left unprocessed. If a section is incomplete, it is revisited\n\n    - **5. Final verification**: Once all sections are summarized, a final check confirms that all titles and subtitles have been correctly processed.\n \n- **Multimodal inputs**: Combined research paper summaries, X-ray descriptions, and discussion topics into a cohesive dataset to allow the model to seamlessly switch between technical, visual and clinical contexts.\n\n- **The Anatomy of a radiologist AI conversation**\n\nEach conversation began with categories and topics, like “Pneumonia Diagnosis on Chest X-rays.” Using summarized papers and X-ray descriptions as references, the AI simulated discussions such as:\n\n- **Dr. A**:\n\n    \"Based on the X-ray, there’s an opacity in the lower lobe consistent with pneumonia. What do you think, Dr. B?\"\n\n- **Dr. B**:\n\n    \"This aligns with the reference paper, 'Imaging in Pneumonia Diagnosis,' which emphasizes lower lobe opacities as early infection markers. Should we explore clinical correlations next?\"\n\n\n### Key lessions for long context application\n\nHere are some of my takeaways if you intend to unlock full potential of long context models\n\n**1. Structure is critical**: By breaking down inputs into categories, topics, and summaries, Gemini can get the clarity needed to perform at its\n\n**2. Summarization saves the day**: Condensed, high quality summaries enables the feeding of more data into the context window while preserving relevances\n\n**Dataset used** [chest-xray-pneumonia](https://www.kaggle.com/datasets/paultimothymooney/chest-xray-pneumonia)\n\n[Youtube video link](https://youtu.be/kxcFZ-Ac6Qw)\n\n![longcontext-usage](/kaggle/input/gemini-longcontext-framework/gemini-longcontext.svg)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"!pip install -q PyMuPDF","metadata":{"execution":{"iopub.status.busy":"2024-11-30T00:40:28.811434Z","iopub.execute_input":"2024-11-30T00:40:28.811883Z","iopub.status.idle":"2024-11-30T00:40:41.604028Z","shell.execute_reply.started":"2024-11-30T00:40:28.811763Z","shell.execute_reply":"2024-11-30T00:40:41.602874Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nfrom pathlib import Path\nfrom tqdm.auto import tqdm\ntqdm.pandas()\nimport glob\n\nimport math\nimport random\nimport pandas as pd\nimport numpy as np\n\nimport cv2\nimport PIL\nimport pydicom\n\nimport fitz\nimport base64\nimport io\nfrom concurrent.futures import ThreadPoolExecutor\nimport requests\n\nimport matplotlib.pyplot as plt\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2024-11-30T00:40:41.606822Z","iopub.execute_input":"2024-11-30T00:40:41.608106Z","iopub.status.idle":"2024-11-30T00:40:42.657179Z","shell.execute_reply.started":"2024-11-30T00:40:41.608067Z","shell.execute_reply":"2024-11-30T00:40:42.656174Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"PIL.Image.open(\"/kaggle/input/gemini-longcontext-use/gemini-1.5-longcontext-use.png\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:40:42.658350Z","iopub.execute_input":"2024-11-30T00:40:42.658790Z","iopub.status.idle":"2024-11-30T00:40:42.726425Z","shell.execute_reply.started":"2024-11-30T00:40:42.658754Z","shell.execute_reply":"2024-11-30T00:40:42.725408Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def download_pdf(url, folder):\n    response = requests.get(url)\n    if response.status_code == 200:\n        file_name = os.path.join(folder, url.split(\"/\")[-1])\n        with open(file_name, 'wb') as file:\n            file.write(response.content)\n        return file_name\n    else:\n        print(f\"Failed to download PDF from {url}\")\n        return None\n\n# Define the function to convert a PDF to a list of base64-encoded PNG images\ndef pdf_to_base64_pngs(pdf_path, quality=75, max_size=(1024, 1024)):\n    # Open the PDF file\n    doc = fitz.open(pdf_path)\n\n    base64_encoded_pngs = []\n\n    # Iterate through each page of the PDF\n    for page_num in range(doc.page_count):\n        # Load the page\n        page = doc.load_page(page_num)\n\n        # Render the page as a PNG image\n        pix = page.get_pixmap(matrix=fitz.Matrix(300/72, 300/72))\n\n        # Convert the pixmap to a PIL Image\n        image = PIL.Image.frombytes(\"RGB\", [pix.width, pix.height], pix.samples)\n\n        # Resize the image if it exceeds the maximum size\n        if image.size[0] > max_size[0] or image.size[1] > max_size[1]:\n            image.thumbnail(max_size, PIL.Image.Resampling.LANCZOS)\n\n        base64_encoded_pngs.append(image)\n\n    # Close the PDF document\n    doc.close()\n\n    return base64_encoded_pngs","metadata":{"execution":{"iopub.status.busy":"2024-11-30T00:40:42.727744Z","iopub.execute_input":"2024-11-30T00:40:42.728057Z","iopub.status.idle":"2024-11-30T00:40:42.736546Z","shell.execute_reply.started":"2024-11-30T00:40:42.728026Z","shell.execute_reply":"2024-11-30T00:40:42.735561Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pneumonia_pdf_urls = [\n    \"https://www.thoracic.org/patients/patient-resources/resources/what-is-pneumonia.pdf\",\n    \"https://www.nhlbi.nih.gov/sites/default/files/publications/what_is_pneumonia.pdf\",\n    \"https://www.vdh.virginia.gov/epidemiology/epidemiology-fact-sheets/pneumonia/?pdf=2252\",\n    \"https://www.pa.gov/content/dam/copapwp-pagov/en/health/documents/topics/documents/diseases-and-conditions/Pneumonia%20.pdf\",\n    \"https://gmch.gov.in/sites/default/files/documents/Pneumonia.pdf\",\n    \"https://www.wsh.nhs.uk/CMS-Documents/Patient-leaflets/Physiotherapy/6356-1-Signs-and-symptoms-of-chest-infections.pdf\",\n    \"https://www.uthsc.edu/pediatrics/clerkship/documents/rti-2.pdf\",\n    \"https://pmc.ncbi.nlm.nih.gov/articles/PMC2601181/pdf/yjbm00554-0079.pdf\",\n    \"https://www.nature.com/articles/s41572-021-00259-0\",\n    \"https://med.stanford.edu/content/dam/sm/bugsanddrugs/documents/clinicalpathways/SHC-Pneumonia-Guideline.pdf\",\n    \"https://jcesom.marshall.edu/media/60719/cap-idsa-pediatrics.pdf\",\n    \"https://www.amcli.it/wp-content/uploads/2019/02/nejmra1714562.pdf\",\n    \"https://www.thelancet.com/journals/langlo/article/PIIS2214-109X(15)00272-7/fulltext\",\n    \"https://ajronline.org/doi/epdf/10.2214/ajr.169.5.9353456\",\n    \"https://bmjopenrespres.bmj.com/content/bmjresp/8/1/e000911.full.pdf\",\n    \"https://link.springer.com/content/pdf/10.1186/cc11201.pdf\",\n    \"https://www.choc.org/wp-content/uploads/2014/12/CHOC_12-27-14_pneumonia.pdf\",\n    \"https://www.lung.org/getmedia/7a41c9f9-d5f9-45ae-966a-9cbc076c9d53/ala-early-warning-signs-one-pager-2021-(4).pdf\",\n    \"https://www.cabrini.com.au/app/uploads/Chest-infection.pdf\",\n]\n\nradiology_pdf_urls = [\n    \"https://aclanthology.org/2024.bionlp-1.55.pdf\",\n    \"https://arxiv.org/pdf/2407.15158\",\n    \"https://openaccess.thecvf.com/content/CVPR2023/papers/Tanida_Interactive_and_Explainable_Region-Guided_Radiology_Report_Generation_CVPR_2023_paper.pdf\",\n    \"https://proceedings.mlr.press/v225/nguyen23a/nguyen23a.pdf\",\n    \"https://med.stanford.edu/content/dam/sm/radiology/documents/about/annualreport/2017-19RadiologyDeptReport-web.pdf\",\n    \"https://www.scienceopen.com/document_file/670c1964-bf8d-4a9e-82b8-90174d59ab83/PubMedCentral/670c1964-bf8d-4a9e-82b8-90174d59ab83.pdf\",\n    \"https://geiselmed.dartmouth.edu/radiology/wp-content/uploads/sites/47/2021/07/The-written-radiology-report.pdf\",\n    \"https://nt-e.nl/assets/blogfiles/How-to-create-a-great-radiology-report.-Hartung-et-al.-2020.pdf\",\n    \"https://med.und.edu/education-training/radiology/_files/docs/xray-film-reading-made-easy.pdf\",\n    \"https://www.acr.org/-/media/acr/files/practice-parameters/communicationdiag.pdf\",\n    \"https://www.acr.org/-/media/ACR/Files/RADS/BI-RADS/Mammography-Reporting.pdf\",\n    \"https://pubs.rsna.org/doi/epdf/10.1148/rg.2020200020\",\n    \"https://www.southafrica-usa.net/homeaffairs/forms/bi806.pdf\",\n    \"https://www.acadrad.org/wp-content/uploads/2018/05/1-Radiographics-Multimedia-Reporting-2017170047.pdf\",\n    \"https://www.aapm.org/pubs/reports/rpt_74.pdf\",\n]\n\nlen(radiology_pdf_urls), len(pneumonia_pdf_urls)","metadata":{"execution":{"iopub.status.busy":"2024-11-30T00:40:42.739934Z","iopub.execute_input":"2024-11-30T00:40:42.740233Z","iopub.status.idle":"2024-11-30T00:40:42.759767Z","shell.execute_reply.started":"2024-11-30T00:40:42.740204Z","shell.execute_reply":"2024-11-30T00:40:42.758755Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pneumonia_folder = \"/kaggle/working/pneumonia\"\nradiology_folder = \"/kaggle/working/radiology\"\n\nos.makedirs(pneumonia_folder, exist_ok=True)\nos.makedirs(radiology_folder, exist_ok=True)\n\nwith ThreadPoolExecutor() as executor:\n    pneumonia_pdf_paths = list(\n        executor.map(download_pdf, pneumonia_pdf_urls, [pneumonia_folder]*len(pneumonia_pdf_urls))\n    )\n    \n    radiology_pdf_paths = list(\n        executor.map(download_pdf, radiology_pdf_urls, [radiology_folder]*len(radiology_pdf_urls))\n    )\n    \n\npneumonia_pdf_paths = [path for path in pneumonia_pdf_paths if path is not None]\nradiology_pdf_paths = [path for path in radiology_pdf_paths if path is not None]","metadata":{"execution":{"iopub.status.busy":"2024-11-30T00:40:42.760953Z","iopub.execute_input":"2024-11-30T00:40:42.761302Z","iopub.status.idle":"2024-11-30T00:41:15.207611Z","shell.execute_reply.started":"2024-11-30T00:40:42.761247Z","shell.execute_reply":"2024-11-30T00:41:15.206704Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class CONFIG:\n    DATA_PATH = Path(\"/kaggle/input/chest-xray-pneumonia/chest_xray\")\n    TRAIN_PATH = DATA_PATH/'train'\n    VALID_PATH = DATA_PATH/'val'\n    TEST_PATH = DATA_PATH/'test'\n    num_images = 50\n    \ncfg = CONFIG()\nos.listdir(cfg.TRAIN_PATH)","metadata":{"execution":{"iopub.status.busy":"2024-11-30T00:41:15.209071Z","iopub.execute_input":"2024-11-30T00:41:15.210137Z","iopub.status.idle":"2024-11-30T00:41:15.222504Z","shell.execute_reply.started":"2024-11-30T00:41:15.210100Z","shell.execute_reply":"2024-11-30T00:41:15.221457Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# os.listdir(cfg.TRAIN_PATH/'NORMAL')\nnormal = sorted(glob.glob(f\"{cfg.TRAIN_PATH}/NORMAL/*.jpeg\"))\npneumonia = sorted(glob.glob(f\"{cfg.TRAIN_PATH}/PNEUMONIA/*.jpeg\"))","metadata":{"execution":{"iopub.status.busy":"2024-11-30T00:41:15.223763Z","iopub.execute_input":"2024-11-30T00:41:15.224075Z","iopub.status.idle":"2024-11-30T00:41:15.404599Z","shell.execute_reply.started":"2024-11-30T00:41:15.224043Z","shell.execute_reply":"2024-11-30T00:41:15.403674Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"FRAMES = []\n\nfor i, path in enumerate(normal[:100]):\n    frame = cv2.imread(path)\n    frame = cv2.resize(frame, (384, 384))\n    frame = PIL.Image.fromarray(frame)\n#     frame.save(f\"chest_{i}.jpeg\")\n    FRAMES.append(frame)","metadata":{"execution":{"iopub.status.busy":"2024-11-30T00:41:15.405752Z","iopub.execute_input":"2024-11-30T00:41:15.406081Z","iopub.status.idle":"2024-11-30T00:41:18.243707Z","shell.execute_reply.started":"2024-11-30T00:41:15.406050Z","shell.execute_reply":"2024-11-30T00:41:18.242548Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"FRAMES[0].save('xray.gif', save_all=True, append_images=FRAMES[1:], duration=360, loop=0)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:41:18.244904Z","iopub.execute_input":"2024-11-30T00:41:18.245230Z","iopub.status.idle":"2024-11-30T00:41:21.062821Z","shell.execute_reply.started":"2024-11-30T00:41:18.245197Z","shell.execute_reply":"2024-11-30T00:41:21.061879Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from IPython.display import Image\nImage(open('/kaggle/working/xray.gif','rb').read())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:41:21.064453Z","iopub.execute_input":"2024-11-30T00:41:21.064835Z","iopub.status.idle":"2024-11-30T00:41:21.420927Z","shell.execute_reply.started":"2024-11-30T00:41:21.064778Z","shell.execute_reply":"2024-11-30T00:41:21.419875Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Gemini Setup","metadata":{}},{"cell_type":"code","source":"from kaggle_secrets import UserSecretsClient\nimport google.generativeai as genai\nfrom google.api_core import retry\nfrom IPython.display import Markdown\n\nuser_secrets = UserSecretsClient()\nsecret_value_0 = user_secrets.get_secret(\"GEMINI_API\")\ngenai.configure(api_key=secret_value_0)\nprint(\"API KEY configured successfully ...\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:41:21.422226Z","iopub.execute_input":"2024-11-30T00:41:21.422572Z","iopub.status.idle":"2024-11-30T00:41:22.805031Z","shell.execute_reply.started":"2024-11-30T00:41:21.422540Z","shell.execute_reply":"2024-11-30T00:41:22.803895Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def get_model(model_name:str = \"gemini-1.5-flash-8b\", temp:float = 0.0, system_instruction:str = None) -> genai.GenerativeModel:\n    model = genai.GenerativeModel(\n        model_name=model_name,\n        generation_config={\n            \"temperature\": temp,\n        },\n        system_instruction=system_instruction\n    )\n    return model","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:41:22.806195Z","iopub.execute_input":"2024-11-30T00:41:22.806668Z","iopub.status.idle":"2024-11-30T00:41:22.812491Z","shell.execute_reply.started":"2024-11-30T00:41:22.806637Z","shell.execute_reply":"2024-11-30T00:41:22.811413Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## PDF Files","metadata":{}},{"cell_type":"code","source":"# pneumonia_pdf_paths\nradiology_pdf_files = [\n    genai.upload_file(p, mime_type='application/pdf') for i, p in enumerate(tqdm(radiology_pdf_paths))\n    if os.path.basename(p).split('.')[-1] == 'pdf'\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:41:22.817036Z","iopub.execute_input":"2024-11-30T00:41:22.817396Z","iopub.status.idle":"2024-11-30T00:41:45.183304Z","shell.execute_reply.started":"2024-11-30T00:41:22.817363Z","shell.execute_reply":"2024-11-30T00:41:45.182287Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# pneumonia_pdf_paths\npneumonia_pdf_files = [\n    genai.upload_file(p, mime_type='application/pdf') for i, p in enumerate(tqdm(pneumonia_pdf_paths))\n    if os.path.basename(p).split('.')[-1] == 'pdf'\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:41:45.184678Z","iopub.execute_input":"2024-11-30T00:41:45.184999Z","iopub.status.idle":"2024-11-30T00:42:02.764493Z","shell.execute_reply.started":"2024-11-30T00:41:45.184968Z","shell.execute_reply":"2024-11-30T00:42:02.763494Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Images","metadata":{}},{"cell_type":"code","source":"%%time\n# cfg.num_images = 10\nn = cfg.num_images\ntrain_normal = random.sample(normal, n)\ntrain_pneumonia = random.sample(pneumonia, n)\n\nnormal_files = [\n    genai.upload_file(p) for i, p in enumerate(train_normal)\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:42:02.765914Z","iopub.execute_input":"2024-11-30T00:42:02.766344Z","iopub.status.idle":"2024-11-30T00:43:51.007007Z","shell.execute_reply.started":"2024-11-30T00:42:02.766296Z","shell.execute_reply":"2024-11-30T00:43:51.005978Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pneumonia_files = [\n    genai.upload_file(p) for i, p in enumerate(train_pneumonia)\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:43:51.008514Z","iopub.execute_input":"2024-11-30T00:43:51.009254Z","iopub.status.idle":"2024-11-30T00:45:37.932031Z","shell.execute_reply.started":"2024-11-30T00:43:51.009204Z","shell.execute_reply":"2024-11-30T00:45:37.930863Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# X-ray description","metadata":{}},{"cell_type":"code","source":"xray_instructions = \"\"\"\n**Normal (No Pneumonia)**:\n- The lungs appear clear and well-inflated with no visible consolidation or infiltrates. The pulmonary vasculature is visible without signs of widening or blurring. No abnormal densities or opacities are observed.\n\n**Early Pneumonia**:\n- The X-ray shows slight opacity in the lower lobe, consistent with early consolidation. There is a subtle increase in lung density, which may suggest the beginning stages of pneumonia. Some peripheral infiltrates are noted, but they are not extensive.\n\n**False Positive Consideration**:\n- In rare cases, certain viral infections can cause mild opacity that could be mistaken for early pneumonia. A careful comparison with the patient's clinical symptoms and history would help avoid this misinterpretation.\n\n**Edge Case**:\n- In this case, the pneumonia signs are minimal, and the early-stage infection could easily be mistaken for a minor viral infection. The absence of significant consolidation makes the diagnosis challenging.\n\"\"\"\n\nimg_prompt = \"\"\"\nProvide a detailed and structured description of the pneumonia cases based on the following x-rays. \nThese descriptions will be used as part of a deep analysis, focusing on interpreting the visual features of the lungs in these images.\n\nThe images are categorized as follows:\n1. **Normal (No Pneumonia)**: These x-rays depict healthy lungs, with no visible signs of pneumonia or abnormality.\n2. **Early Pneumonia**: These x-rays show the early stages of pneumonia, where early signs of infection or inflammation may be present.\n\nFor each image:\n- **Describe the visual features** of the lungs, focusing on abnormalities that may indicate early pneumonia.\n- **Highlight differences** between the normal and early pneumonia x-rays, including any signs such as consolidation, opacity, infiltrates, or other lung changes.\n- **Provide an insightful analysis** of how early pneumonia can be detected visually, and what specific features in the x-rays stand out as indicators of infection.\n- **Examine edge cases** where the pneumonia may be difficult to detect due to early, subtle changes in the lungs, or when it mimics other conditions.\n- **Identify false positives and negatives**: Discuss scenarios where healthy lungs may appear abnormal (false positive) or where early pneumonia may be missed or misinterpreted as normal (false negative). Describe why these mistakes could happen and what to look for to avoid them.\n- Ensure that the descriptions are thorough and precise, as they will be used to inform further analysis and decision-making.\n\n### Normal (No Pneumonia) X-rays:\n{normal_images}\n\n### Early Pneumonia X-rays:\n{pneumonia_images}\n\"\"\"\n\nmodel = get_model(model_name=\"gemini-1.5-flash-latest\", system_instruction=xray_instructions)\nxray_response = model.generate_content(\n    [img_prompt, \"\\n\\nNormal x-rays:\\n\\n\", *normal_files, \"\\n\\nPneumonia x-rays:\\n\\n\", *pneumonia_files]\n)\n\nxray_response = xray_response.text\nMarkdown(xray_response)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:45:37.933399Z","iopub.execute_input":"2024-11-30T00:45:37.933735Z","iopub.status.idle":"2024-11-30T00:45:48.553587Z","shell.execute_reply.started":"2024-11-30T00:45:37.933703Z","shell.execute_reply":"2024-11-30T00:45:48.552527Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Summarizer","metadata":{}},{"cell_type":"code","source":"prompt = \"\"\"\nYou are tasked with analyzing a PDF document that has been divided into individual pages. Each page contains valuable information, including titles, subtitles, figures, and tables. Your goal is to process **every single chapter, title, and subtitle**, ensuring no content is omitted except for the references section. The model must perform a self-check to verify that all identified sections are processed.\n\n### Step-by-Step Instructions:\n\n1. **Identify the Paper Title**:\n   - Begin by stating the **Paper Title** to provide context for the document being analyzed.\n\n2. **Identify Titles and Subtitles**:\n   - After the paper title, extract and list all titles, subtitles, and headings from the entire document.\n   - Organize them hierarchically to reflect the document's structure.\n   - Clearly state that this is the \"List of Titles and Subtitles\" before starting the summaries.\n\n3. **Process Each Title or Subtitle**:\n   - For each identified title or subtitle:\n     - **Section Title (or Subtitle)**: State the title or main heading.\n     - **Detailed Summary**: Provide a detailed summary of the section, including:\n       - The main ideas or findings.\n       - Supporting evidence, examples, or data points.\n       - Implications, applications, or conclusions.\n       - Connections to broader themes, concepts, or other sections.\n     - **Figures and Tables**: For each figure or table in the section:\n       - Describe the figure or table content (e.g., \"Figure 1: A radiograph showing features of alveolar infiltrates\").\n       - Summarize key insights or conclusions derived from the figure or table.\n       - Explain how it supports the section's findings or arguments.\n\n4. **Self-Check Mechanism**:\n   - After summarizing each title or subtitle, compare it against the original \"List of Titles and Subtitles\".\n   - Clearly indicate whether the task for that title or subtitle has been completed.\n   - If any titles or subtitles remain incomplete, revisit them immediately and complete their summaries.\n   - Once all titles and subtitles in the list are processed, explicitly confirm: \"All sections have been completed.\"\n\n5. **Output Structure**:\n   - Begin with the **Paper Title** followed by the \"List of Titles and Subtitles.\"\n   - For each title or subtitle in the list, include its detailed summary and any associated figures or tables.\n   - Ensure all sections are fully processed before concluding.\n\n### Example of Expected Output:\n\n#### Paper Title: The Impact of Radiology in Diagnosing Pneumonia\n#### List of Titles and Subtitles:\n1. Chapter One: Introduction\n2. Chapter Two: Chest\n   - 2.1 Overview of Chest Imaging\n   - 2.2 Techniques for Radiograph Analysis\n3. Chapter Three: Abdomen\n\n#### Detailed Summaries:\n\n- **Chapter One: Introduction**\n  - **Detailed Summary**:\n    - Key idea: The introduction provides an overview of the document’s purpose and sets the context for subsequent sections.\n    - Supporting evidence: Discusses trends in radiology and recent advancements in imaging technology.\n    - Implications: Sets the foundation for understanding the challenges in diagnostic imaging.\n\n- **Chapter Two: Chest**\n  - **Detailed Summary**:\n    - Key idea: This chapter outlines a systematic approach to interpreting chest radiographs.\n    - Supporting evidence:\n      - Highlights differences between interstitial and alveolar infiltrates.\n      - Includes Figure 4: A photomicrograph of normal lung tissue, illustrating typical cellular structures.\n    - Figures and Tables:\n      - **Figure 4**: Photomicrograph of normal lung tissue. **Key takeaway**: Serves as a baseline for comparison with pathological findings.\n\n- **Self-Check**:\n  - \"Chapter One: Introduction\": Completed ✅\n  - \"Chapter Two: Chest\": Completed ✅\n  - Remaining sections to process: [\"Chapter Three: Abdomen\"]\n\n- **Chapter Three: Abdomen**\n  - [Detailed Summary and Figures/Tables]\n\n#### Final Self-Check:\n- All sections processed: ✅ Yes / ❌ No.\n- If \"No,\" process the remaining sections.\n\n### Important Notes:\n- Perform the self-check diligently to ensure no title or subtitle is left unprocessed.\n- Summaries must be comprehensive, insightful, and professional.\n- Maintain consistent detail and depth across all sections.\n\"\"\"\n\nmodel = get_model(model_name=\"gemini-1.5-flash-8b\")\nresponse = model.generate_content(\n    [radiology_pdf_files[2], prompt]\n)\n\nMarkdown(response.text)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:45:48.555146Z","iopub.execute_input":"2024-11-30T00:45:48.555589Z","iopub.status.idle":"2024-11-30T00:45:58.212514Z","shell.execute_reply.started":"2024-11-30T00:45:48.555543Z","shell.execute_reply":"2024-11-30T00:45:58.211310Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import asyncio\nfrom typing import List\n\nasync def summarize_pdf(pdf_file, prompt):\n    model = genai.GenerativeModel(\n        model_name=\"gemini-1.5-flash-8b\",\n    )\n    \n    try:\n        response = await asyncio.to_thread(model.generate_content, [pdf_file, prompt])\n        \n        return response.text\n    except Exception as e:\n        print(f\"Error summarizing document {pdf_file.display_name}: {e}\")\n        return None\n\nasync def summarize_documents(radiology_refs, pneumonia_refs, prompt):\n    summaries = []\n    max_files = max(len(radiology_refs), len(pneumonia_refs))\n    for i in range(max_files):\n        print(\"#\"*(i+1), end='')\n        if i < len(radiology_refs):\n            radiology_pdf = radiology_refs[i]\n            radiology_summary = await summarize_pdf(radiology_pdf, prompt)\n            if radiology_summary:\n                summaries.append(f\"Radiology Document {i+1} Summary:\\n{radiology_summary}\\n\")\n\n        if i < len(pneumonia_refs):\n            pneumonia_pdf = pneumonia_refs[i]\n            pneumonia_summary = await summarize_pdf(pneumonia_pdf, prompt)\n            if pneumonia_summary:\n                summaries.append(f\"Pneumonia Document {i+1} Summary:\\n{pneumonia_summary}\\n\")\n    return summaries\n    \n# summarize_pdf(model, radiology_pdf_files[0], prompt)\ndetailed_summary = await summarize_documents(radiology_pdf_files, pneumonia_pdf_files, prompt=prompt)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:45:58.213777Z","iopub.execute_input":"2024-11-30T00:45:58.214096Z","iopub.status.idle":"2024-11-30T00:50:02.931210Z","shell.execute_reply.started":"2024-11-30T00:45:58.214065Z","shell.execute_reply":"2024-11-30T00:50:02.930166Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Generate topics","metadata":{}},{"cell_type":"code","source":"topic_prompt = \"\"\"\nYou are tasked with generating a list of discussion topics based on the radiology and pneumonia references provided below. Your goal is to identify all relevant topics that could guide a detailed and insightful conversation. \n\n### Guidelines for Generating the Topic List:\n\n1. **Review the References**:\n   - Examine the radiology and pneumonia references carefully.\n   - Identify major themes, subtopics, and key insights from each document.\n\n2. **Categorize Topics**:\n   - Group related topics under broader categories, such as:\n     - Imaging Techniques and Findings\n     - Differential Diagnoses\n     - Technological Advancements\n     - Clinical Implications and Treatment Options\n\n3. **Structure the List**:\n   - For each category, provide a detailed list of topics or questions. \n   - Ensure that topics are specific and cover all critical aspects of the references.\n   - Include topics that connect radiology and pneumonia insights (e.g., imaging findings specific to pneumonia cases).\n   - When referencing a paper, **use the exact Paper Title** from the summary of the reference.\n\n4. **Ensure Completeness**:\n   - Cross-reference all documents to ensure that no major topics are omitted.\n   - Topics should address both technical and clinical perspectives.\n   - When referring to any paper, include the **Paper Title** to ensure proper citation.\n\n### Format for the Output:\n\n- **Category 1: [Broad Theme]**\n  - Topic 1: [Specific topic or question].\n  - Topic 2: [Specific topic or question].\n  ...\n\n- **Category 2: [Broad Theme]**\n  - Topic 1: [Specific topic or question].\n  - Topic 2: [Specific topic or question].\n  ...\n\nRepeat for all categories and topics.\n\n### References:\n\n#### Radiology References:\n- **Paper Title: Radiological Insights into Lung Diseases**\n- **Paper Title: Advanced Imaging Techniques in Radiology**\n- **Paper Title: The Role of CT in Diagnosing Pneumonia**\n... (Add more as needed)\n\n#### Pneumonia References:\n- **Paper Title: Clinical Approaches to Pneumonia Diagnosis**\n- **Paper Title: Understanding Pneumonia Pathophysiology**\n- **Paper Title: Pneumonia in the Elderly: Diagnosis and Treatment**\n... (Add more as needed)\n\n**Important Notes**:\n- Do not generate the actual discussion content, only the topic list.\n- Ensure the topics are detailed, well-organized, and cover all key aspects of the references.\n- **Always cite papers by their Paper Title** as they appear in the reference summaries.\n\"\"\"\n\nmodel = get_model(model_name=\"gemini-1.5-flash-8b\")\n\nresp2 = model.generate_content(\n    [topic_prompt, \"\\n\\n\", *detailed_summary]\n)\n\nMarkdown(resp2.text)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:50:02.932731Z","iopub.execute_input":"2024-11-30T00:50:02.933145Z","iopub.status.idle":"2024-11-30T00:50:13.609466Z","shell.execute_reply.started":"2024-11-30T00:50:02.933101Z","shell.execute_reply":"2024-11-30T00:50:13.608338Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import re\n\n# Simulated parse_topics function\ndef parse_topics(text):\n    # Split the text into categories and topics based on the structure\n    parsed_output = {}\n    \n    # Regular expression to match the category and its content\n    category_pattern = r\"\\*\\*Category (\\d+): (.*?)\\*\\*\\s*(.*?)\\s*(?=\\*\\*Category|\\Z)\"\n    \n    # Match each category and its content\n    categories = re.findall(category_pattern, text, re.DOTALL)\n    \n    for category_num, category_name, category_content in categories:\n        topics = []\n        \n        # Regex to find individual topics within the category\n        topic_pattern = r\"- Topic (\\d+): (.*?)\\n\"\n        topics_found = re.findall(topic_pattern, category_content)\n        \n        for topic_num, topic_details in topics_found:\n            topics.append({\n                'topic_number': topic_num,\n                'details': topic_details\n            })\n        \n        parsed_output[category_name] = {\n            'category_number': category_num,\n            'name': category_name,\n            'topics': topics\n        }\n    \n    return parsed_output\n\ntext = resp2.text\n# Parsing the text\nparsed_output = parse_topics(text)\n\n# Generating Category Instructions\nCATEGORY_INSTRUCTIONS = \"You are two radiologists, Dr. A and Dr. B, discussing various aspects of pneumonia and its diagnosis based on imaging techniques and clinical data. The discussion should be detailed, technical, and incorporate references to the provided documents. Ensure the conversation flows naturally, and keep it relevant to the topics in the categories.\\n\\nHere are the categories and topics you will discuss:\\n\"\n\n# Loop through the parsed data to create the instruction text\nfor category, data in parsed_output.items():\n    CATEGORY_INSTRUCTIONS += f\"\\n1. **Category: {data['name']}**\\n\"\n    for topic in data['topics']:\n        CATEGORY_INSTRUCTIONS += f\"    - Topic {topic['topic_number']}: {topic['details']}\\n\"","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:50:13.610593Z","iopub.execute_input":"2024-11-30T00:50:13.610919Z","iopub.status.idle":"2024-11-30T00:50:13.620590Z","shell.execute_reply.started":"2024-11-30T00:50:13.610887Z","shell.execute_reply":"2024-11-30T00:50:13.619455Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Simulate discussion\n\nThe input prompt is divided into two sections. The initial prompt contains only the `detailed summary` and `x-ray description`, after this initial phase, then prompt will also contain the past discussion.","metadata":{}},{"cell_type":"code","source":"model = get_model(model_name=\"gemini-1.5-flash-8b\", temp=2.0,system_instruction=CATEGORY_INSTRUCTIONS)\n\n# Simulating the discussion for each category\nlast_response = \"\"\nfor category, data in parsed_output.items():\n    print(\"-*\" * 42)\n    \n    # Creating the category prompt (introductory message for the first category)\n    if category == list(parsed_output.keys())[0]:\n        cat_prompt = f\"\"\"\n        Start the discussion by introducing the topics of {category}: {data['name']}.\n        The radiologists, Dr. A and Dr. B, will now discuss the topics outlined below. The conversation should flow naturally and be detailed.\n        \"\"\"\n    else:\n        # For subsequent categories, no need for an introductory message; just continue the discussion\n        cat_prompt = f\"\"\"\n        Dr. A and Dr. B continue discussing the next category: {data['name']}. \n        Previously: {last_response}. \n        The conversation should continue seamlessly, maintaining the natural flow.\n        \"\"\"\n    \n    # Using the model to generate the discussion for each category\n    resp3 = model.generate_content(\n        [cat_prompt, \"\\n\\nResearch papers\\n\\n\", *detailed_summary, \"\\n\\nX-ray description\\n\\n\", xray_response]  # Assuming 'output' holds prior discussion content\n    )\n\n    # Update the response to dynamically include paper titles as citations\n    last_response += resp3.text\n    \n    # Print the generated response for each category\n    print(resp3.text)\n    print()  # Add space between discussions for clarity`","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:50:13.621987Z","iopub.execute_input":"2024-11-30T00:50:13.622326Z","iopub.status.idle":"2024-11-30T00:51:32.836920Z","shell.execute_reply.started":"2024-11-30T00:50:13.622258Z","shell.execute_reply":"2024-11-30T00:51:32.835676Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Number of words","metadata":{}},{"cell_type":"code","source":"print(f\"Initial input size: {len(' '.join(detailed_summary)) + len(xray_response)}\")\nprint(f\"Final input size: {len(last_response) + len(' '.join(detailed_summary)) + len(xray_response)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-30T00:53:42.406499Z","iopub.execute_input":"2024-11-30T00:53:42.406976Z","iopub.status.idle":"2024-11-30T00:53:42.413618Z","shell.execute_reply.started":"2024-11-30T00:53:42.406929Z","shell.execute_reply":"2024-11-30T00:53:42.412424Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}