TIL: Running prompts against images, PDFs, audio and video with Google Gemini

23rd October 2024

TIL Running prompts against images, PDFs, audio and video with Google Gemini — I'm still working towards adding multi-modal support to my [LLM](https://llm.datasette.io/) tool. In the meantime, here are notes on running prompts against images and PDFs and audio and video files from the command-line using the [Google Gemini](https://ai.google.dev/gemini-api) family of models.

Posted 23rd October 2024 at 5:49 pm

Simon Willison’s Weblog

Recent articles

Monthly briefing